This is really insane to me. There's nothing practical about open-source models yet that makes them even remotely comparable to closed frontier models. All the hype around GLM, Qwen, now Kimi.... Are people really this naive that they believe these reports or, more worringly, are people NOT using these models and seeing the HUGE gap that still exists? Take a task, any medium-sized task, decently scoped that you'd tru…
> Take a task, any medium-sized task, decently scoped that you'd trust to give to Sonnet to finish without a hitch. Now give it to ANY open-source frontier model and watch them struggle and go in circles while failing tool calls and randomly assuming things. Claude used to be much worse than it is now, just as bad the open weights models are. And the open weights were worse. The labs will also try to keep the lead, b…
The biggest moat of these giant labs and models is increasingly shifting towards deployment capabilities and (debatably) having better (proprietary) harnesses.
The models themselves can be impressive on benchmarks, but unless they can be served reliably to customers either at scale, hosted somewhere, or even on edge with predictable latency and memory usage, then frontier will always be leading.