Disclaimer: I did not test this yet. I don't want to make big generalizations. But one thing I noticed with chinese models, especially Kimi, is that it does very well on benchmarks, but fails on vibe testing. It feels a little bit over-fitting to the benchmark and less to the use cases. I hope it's not the same here.
This used to happen with bench marks on phones, manufacturers would tweak android so benchmarks ran faster. I guess that’s kinda how it is for any system that’s trained to do well on benchmarks, it does well but rubbish at everything else.
DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
61–70 of 485 posts
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#62 % curl https://api.deepseek.com/models \
-H "Authorization: Bearer ${DEEPSEEK_API_KEY}"
{"object":"list","data":[{"id":"deepseek-chat","object":"model","owned_by":"deepseek"},{"id":"deepseek-reasoner","object":"model","owned_by":"deepseek"}]}Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#63Disclaimer: I did not test this yet. I don't want to make big generalizations. But one thing I noticed with chinese models, especially Kimi, is that it does very well on benchmarks, but fails on vibe testing. It feels a little bit over-fitting to the benchmark and less to the use cases. I hope it's not the same here.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#64I genuinely do not understand the evaluations of the US AI industry. The chinese models are so close and far cheaper
It's all about the hardware and infrastructure. If you check OpenRouter, no provider offers a SOTA chinese model matching the speed of Claude, GPT or Gemini. The chinese models may benchmark close on paper, but real-world deployment is different. So you either buy your own hardware in order to run a chinese model at 150-200tps or give up an use one of the Big 3. The US labs aren't just selling models, they're selling…
Opus 4.5 = ~60-80tps https://openrouter.ai/anthropic/claude-opus-4.5
Kimi-k2-think = ~60-180tps https://openrouter.ai/moonshotai/kimi-k2-thinking
Deepseek-v3.2 = ~30-110tps (only 2 providers rn) https://openrouter.ai/deepseek/deepseek-v3.2
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#65Earlier quoted context omitted.
> Valuation is not based on what they have done but what they might do Exactly what I’m thinking. Chinese models catching rapidly. Soon to be on-par with the big dogs.
Even if they do continue to lag behind they are a good bet against monopolisation by proprietary vendors.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#66Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#67Well props to them for continuing to improve, winning on cost-effectiveness, and continuing to publicly share their improvements. Hard not to root for them as a force to prevent an AI corporate monopoly/duopoly.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#68Disclaimer: I did not test this yet. I don't want to make big generalizations. But one thing I noticed with chinese models, especially Kimi, is that it does very well on benchmarks, but fails on vibe testing. It feels a little bit over-fitting to the benchmark and less to the use cases. I hope it's not the same here.
What is "Vibe testing"?
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#69Earlier quoted context omitted.
Valuation is not based on what they have done but what they might do, I agree tho it's investment made with very little insight into Chinese research. I guess it's counting on deepseek being banned and all computers in America refusing to run open software by the year 2030 /snark
> I guess it's counting on deepseek being banned And the people making the bets are in a position to make sure the banning happens. The US government system being what it is. Not that our leaders need any incentive to ban Chinese tech in this space. Just pointing out that it's not necessarily a "bet". "Bet" imply you don't know the outcome and you have no influence over the outcome. Even "investment" implies you don'…
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#70Earlier quoted context omitted.
Two aspects to consider: 1. Chinese models typically focus on text. US and EU models also bear the cross of handling image, often voice and video. Supporting all those is additional training costs not spent on further reasoning, tying one hand in your back to be more generally useful. 2. The gap seems small, because so many benchmarks get saturated so fast. But towards the top, every 1% increase in benchmarks is sign…
Nothing you said helps with the issue of valuation. Yes, the US models may be better by a few percentage points, but how can they justify being so costly, both operationally as well as in investment costs? Over the long run, this is a business and you don't make money being the first, you have to be more profitable overall.
But even assuming that we're on the right path for that (which we may not be) and assuming that nothing intervenes to stop it (which it might), there may be only one winner, and that winner may not have even entered the game yet.