Disclaimer: I did not test this yet. I don't want to make big generalizations. But one thing I noticed with chinese models, especially Kimi, is that it does very well on benchmarks, but fails on vibe testing. It feels a little bit over-fitting to the benchmark and less to the use cases. I hope it's not the same here.
What is "Vibe testing"?
DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
71–80 of 485 posts
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#72Earlier quoted context omitted.
It's all about the hardware and infrastructure. If you check OpenRouter, no provider offers a SOTA chinese model matching the speed of Claude, GPT or Gemini. The chinese models may benchmark close on paper, but real-world deployment is different. So you either buy your own hardware in order to run a chinese model at 150-200tps or give up an use one of the Big 3. The US labs aren't just selling models, they're selling…
Gemini 3 = ~70tps https://openrouter.ai/google/gemini-3-pro-preview Opus 4.5 = ~60-80tps https://openrouter.ai/anthropic/claude-opus-4.5 Kimi-k2-think = ~60-180tps https://openrouter.ai/moonshotai/kimi-k2-thinking Deepseek-v3.2 = ~30-110tps (only 2 providers rn) https://openrouter.ai/deepseek/deepseek-v3.2
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#73which version are you?
我是DeepSeek最新版本模型! 如果你想了解具体的版本号信息,我建议你:
查看官方文档 - DeepSeek官网和文档会有最准确的版本信息
关注官方公告 - 版本更新通常会在官方渠道公布
查看应用商店/网页版 - 使用界面通常会显示当前版本
我具备DeepSeek的所有最新功能特性,包括: 强大的对话和推理能力
128K上下文长度
文件上传处理(图像、文档等)
联网搜索功能(需手动开启)
完全免费使用
如果你需要知道精确的版本号用于技术对接或其他特定用途,最好直接查阅官方技术文档,那里会有最准确和详细的技术规格说明。有什么其他问题我可以帮你解答吗?
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#74At first, I thought they had gone the route of implementing yet another chat format that can handle more dynamic conversations like that, instead of just using Harmony, but looking at the syntax, doesn't it look exactly like Harmony? That's a good thing, don't get me wrong, but why not mention straight up that they've implemented Harmony, so people can already understand up front that it's compatible with whatever parsing we're using for GPT-OSS?
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#75Disclaimer: I did not test this yet. I don't want to make big generalizations. But one thing I noticed with chinese models, especially Kimi, is that it does very well on benchmarks, but fails on vibe testing. It feels a little bit over-fitting to the benchmark and less to the use cases. I hope it's not the same here.
If it had vision and was better on long context I'd use it so much more.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#76I genuinely do not understand the evaluations of the US AI industry. The chinese models are so close and far cheaper
With caching the expensive US models end up being like 2x the price (e.g sonnet) and often much cheaper (e.g gpt-5 mini)
If they start caching then US companies will be completely out priced.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#77Earlier quoted context omitted.
Nothing you said helps with the issue of valuation. Yes, the US models may be better by a few percentage points, but how can they justify being so costly, both operationally as well as in investment costs? Over the long run, this is a business and you don't make money being the first, you have to be more profitable overall.
I think the investment race here is an "all-pay auction"*. Lots of investors have looked at the ultimate prize — basically winning something larger than the entire present world economy forever — and think "yes". But even assuming that we're on the right path for that (which we may not be) and assuming that nothing intervenes to stop it (which it might), there may be only one winner, and that winner may not have even…
This is what people like Altman want investors to believe. It seems like any other snake oil scam because it doesn't match reality of what he delivers.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#78I genuinely do not understand the evaluations of the US AI industry. The chinese models are so close and far cheaper
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#79Earlier quoted context omitted.
Gemini 3 = ~70tps https://openrouter.ai/google/gemini-3-pro-preview Opus 4.5 = ~60-80tps https://openrouter.ai/anthropic/claude-opus-4.5 Kimi-k2-think = ~60-180tps https://openrouter.ai/moonshotai/kimi-k2-thinking Deepseek-v3.2 = ~30-110tps (only 2 providers rn) https://openrouter.ai/deepseek/deepseek-v3.2
It doesn't work like that. You need to actually use the model and then go to /activity to see the actual speed. I constantly get 150-200tps from the Big 3 while other providers barely hit 50tps even though they advertise much higher speeds. GLM 4.6 via Cerebras is the only one faster than the closed source models at over 600tps.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#80Earlier quoted context omitted.
Two aspects to consider: 1. Chinese models typically focus on text. US and EU models also bear the cross of handling image, often voice and video. Supporting all those is additional training costs not spent on further reasoning, tying one hand in your back to be more generally useful. 2. The gap seems small, because so many benchmarks get saturated so fast. But towards the top, every 1% increase in benchmarks is sign…
forgive me for bringing politics into it, are chinese LLM more prone to censorship bias than US ones ?