DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
1–10 of 485 posts
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#2Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#3Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#4Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#5Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#6Benchmarks are super impressive, as usual. Interesting to note in table 3 of the paper (p. 15), DS-Speciale is 1st or 2nd in accuracy in all tests, but has much higher token output (50% more, or 3.5x vs gemini 3 in the codeforces test!).
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#7I genuinely do not understand the evaluations of the US AI industry. The chinese models are so close and far cheaper
The nature of the race may change as yet though, and I am unsure if the devil is in the details, as in very specific edge cases that will work only with frontier models ?
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#8I genuinely do not understand the evaluations of the US AI industry. The chinese models are so close and far cheaper
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#9Frontier models are far exceeding even the most hardcore consumer hobbyist requirements. This is even further
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#10I genuinely do not understand the evaluations of the US AI industry. The chinese models are so close and far cheaper
The US labs aren't just selling models, they're selling globally distributed, low-latency infrastructure at massive scale. That's what justifies the valuation gap.
Edit: It looks like Cerebras is offering a very fast GLM 4.6