Earlier quoted context omitted.
> The Chinese models are not good enough for anything other than pair programming, which is just a very last-gen way of using agents. This matches my experience with DeepSeek V4 Pro at Max reasoning, the preview version of the model kept regularly messing things up. About 30-60% of additional time to fix the output was needed. On similar tasks, GLM 5.2 at Max reasoning screwed up maybe 20-30% of the time, while it st…
> This matches my experience with DeepSeek V4 Pro at Max reasoning, That was ages ago (in LLM release timelines). DeepSeek V4 Flash beats it now and a lot cheaper. > On similar tasks, GLM 5.2 at Max reasoning screwed up maybe 20-30% of the time, GLM 5.3 bridges this gap. > I'd say as Chinese models get better, whatever moat Anthropic and OpenAI have dissipates. Their moat, especially OpenAI is funding and hardware re…
I’m sure the next models will only get better, when they’re released. Also super curious about what Moonshot will achieve and the full DeepSeek V4 Pro release!
> Their moat, especially OpenAI is funding and hardware resources. They gain train models 10x as large and also serve at large scale. That's it.
I’ve seen how much slower Kimi K3 can be and that part seems correct, their own GPU production still has ways to go and export restrictions definitely limit what they can do.
Not sure about the size part, if Kimi K3 achieves SOTA performance at 2.8T parameters, western models being >2x that size would be insanely bad in regards to efficiency. I bet they’re all within the same order of magnitude and below 10T and won’t really have a reason to go even that high for the foreseeable future.
As investors will start squeezing them for profitability, I suspect focusing more on efficiency will be commonplace.