Earlier quoted context omitted.
Why would someone hold the opposite opinion? It’s much more likely that our current approach to large language models for general use will eventually show diminishing improvements (even if you think it hasn’t yet), than the opposite situation where valuable improvements can be made forever. The threat of distillation and efficiency gains from competitors mean that providing value at the top end of the market is an ex…
Sorry, this is ridiculous. OpenAI said that they have a step change in model performance. They proved it by solving a Millenium problem. HF incidents are public and vetted. If you still think there's a stall despite all evidence pointing to opposite, I don't know what to say..
Are these real step changes - big picture wise, or refinements in RL/agentic orchestration/"taste" and advancements due to bigger models and hardware technology/capacity scaling? If they not, does this tactic - and hardware improvements - continue to scale non-linearly like they need to?
It is clear that whatever does change in each model increment has resulted in meaningfully better end user capabilities (as well as regressions in some areas, honestly), but that doesn't prove anything. I'm not sure what I personally believe, but stating with your full chest that a stall is ridiculous ignores a lot of potential evidence to the contrary.
Stupid example: Astra. Its main improvements are: much much better computer use and 3d modeling capabilities; better subagent orchestration; better and more reliable tool use; slightly worse coding.
This looks to me, from a feature perspective, to be an incremental improvement across several functional areas, plus new features which are unquestionably excellent but are probably the result of RL focus, not magic.
Step change? Ehhhh depends on how you squint. But how many more iterations of this do we have? Are we going to squeeze quintillion parameter transformers into GPUs?