Earlier quoted context omitted.
there's no reason to believe that performance will continue to scale with compute, though. that's why there's a rout. more simply, if you assume maximum performance with the current LLM/transformer architecture is say, twice as good as what humanity is capable of now, then that would mean that you're approaching 50%+ performance with orders of magnitude less compute. there's just no way you could justify the amount o…
Wait no, there is actually PLENTY of evidence that performance continues to scale with more compute. The entire point of the o3 announcement and benchmark results of throwing a million bucks of test time compute at ARC-AGI is that the ceiling is really really high. We have 3 verified scaling laws of pre-training corpus size, parameter count, and test time compute. More efficiency is fantastic progress, but we will al…
would love to see evidence to the contrary. my assertion comes from seeing claude, gemini and o1.
if anything I feel performance is more of a function of the quality of data than anything else.