Does any expert in the field know whether it is really the case that this intelligence we are seeing with frontier models is an "emerging" phenomena, only coming up when the architecture is scaled? Like isn't it weird that the 1 million parameter model with the same architecture can't solve basic puzzles but suddenly the 1 trillion parameter can conjure up counter-examples for the Jacobian conjecture? It's unintuitiv…
Emergent phenomena = pass/fail grading of multi part problems. If you're just gonna count it as a fail you're not gonna see the progress until everything works, even if every single subproblem improves linearly.