Well props to them for continuing to improve, winning on cost-effectiveness, and continuing to publicly share their improvements. Hard not to root for them as a force to prevent an AI corporate monopoly/duopoly.
DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
31–40 of 485 posts
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#32I genuinely do not understand the evaluations of the US AI industry. The chinese models are so close and far cheaper
It's all about the hardware and infrastructure. If you check OpenRouter, no provider offers a SOTA chinese model matching the speed of Claude, GPT or Gemini. The chinese models may benchmark close on paper, but real-world deployment is different. So you either buy your own hardware in order to run a chinese model at 150-200tps or give up an use one of the Big 3. The US labs aren't just selling models, they're selling…
It'll probably be a few years before all that stuff becomes as smooth as people need, but OAI and Anthropic are already doing a good job on that front.
Each new Chinese model requires a lot of testing and bespoke conformance to every task you want to use it for. There's a lot of activity and shared prompt engineering, and some really competent people doing things out in the open, but it's generally going to take a lot more expert work getting the new Chinese models up to snuff than working with the big US labs. Their product and testing teams do a lot of valuable work.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#33I genuinely do not understand the evaluations of the US AI industry. The chinese models are so close and far cheaper
There is a great deal of orientalism --- it is genuinely unthinkable to a lot of American tech dullards that the Chinese could be better at anything requiring what they think of as "intelligence." Aren't they Communist? Backward? Don't they eat weird stuff at wet markets? It reminds me, in an encouraging way, of the way that German military planners regarded the Soviet Union in the lead-up to Operation Barbarossa. Th…
Ideology played a role, but the data they worked with, was the finnish war, that was disastrous for the sowjet side. Hitler later famously said, it was all a intentionally distraction to make them believe the sowjet army was worth nothing. (Real reasons were more complex, like previous purging).
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#34It's awesome that stuff like this is open source, but even if you have a basement rig with 4 NVIDIA GeForce RTX 5090 graphic cards ($15-20k machine), can it even run with any reasonable context window that isn't like a crawling 10/tps? Frontier models are far exceeding even the most hardcore consumer hobbyist requirements. This is even further
Home rigs like that are no longer cost effective. You're better off buying an rtx pro 6000 outright. This holds both for the sticker price, the supporting hardware price, the electricity cost to run it and cooling the room that you use it in.
https://www.youtube.com/watch?v=zwHqO1mnMsA
I wonder how well the aftermarket memory surgery business on consumer GPUs is doing.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#35what is the ballpark vram / gpu requirement to run this ?
For just the model itself: 4 x params at F32, 2 x params at F16/BF16, or 1 x params at F8, e.g. 685GB at F8. It will be smaller for quantizations, but I'm not sure how to estimate those. For a Mixture of Experts (MoE) model you only need to have the memory size of a given expert. There will be some swapping out as it figures out which expert to use, or to change expert, but once that expert is loaded it won't be swap…
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#36I genuinely do not understand the evaluations of the US AI industry. The chinese models are so close and far cheaper
It's all about the hardware and infrastructure. If you check OpenRouter, no provider offers a SOTA chinese model matching the speed of Claude, GPT or Gemini. The chinese models may benchmark close on paper, but real-world deployment is different. So you either buy your own hardware in order to run a chinese model at 150-200tps or give up an use one of the Big 3. The US labs aren't just selling models, they're selling…
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#37I genuinely do not understand the evaluations of the US AI industry. The chinese models are so close and far cheaper
There is a great deal of orientalism --- it is genuinely unthinkable to a lot of American tech dullards that the Chinese could be better at anything requiring what they think of as "intelligence." Aren't they Communist? Backward? Don't they eat weird stuff at wet markets? It reminds me, in an encouraging way, of the way that German military planners regarded the Soviet Union in the lead-up to Operation Barbarossa. Th…
Though, because Stalin had decimated the red army leadership (including most of the veteran officer who had Russian civil war experience) during the Moscow trials purges, the German almost succeeded.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#38It's awesome that stuff like this is open source, but even if you have a basement rig with 4 NVIDIA GeForce RTX 5090 graphic cards ($15-20k machine), can it even run with any reasonable context window that isn't like a crawling 10/tps? Frontier models are far exceeding even the most hardcore consumer hobbyist requirements. This is even further
Home rigs like that are no longer cost effective. You're better off buying an rtx pro 6000 outright. This holds both for the sticker price, the supporting hardware price, the electricity cost to run it and cooling the room that you use it in.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#39I genuinely do not understand the evaluations of the US AI industry. The chinese models are so close and far cheaper
There is a great deal of orientalism --- it is genuinely unthinkable to a lot of American tech dullards that the Chinese could be better at anything requiring what they think of as "intelligence." Aren't they Communist? Backward? Don't they eat weird stuff at wet markets? It reminds me, in an encouraging way, of the way that German military planners regarded the Soviet Union in the lead-up to Operation Barbarossa. Th…
Stalin just finished purging his entire officer corps, which is not a good omen for war, and the USSR failed miserably against the Finnish who were not the strongest of nations, while Germany just steamrolled France, a country that was much more impressive in WW1 than the Russians (who collapsed against Germany)
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#40Earlier quoted context omitted.
There is a great deal of orientalism --- it is genuinely unthinkable to a lot of American tech dullards that the Chinese could be better at anything requiring what they think of as "intelligence." Aren't they Communist? Backward? Don't they eat weird stuff at wet markets? It reminds me, in an encouraging way, of the way that German military planners regarded the Soviet Union in the lead-up to Operation Barbarossa. Th…
Back when deepseek came out and people were tripping over themselves shouting it was so much better than what was out there, it just wasn’t good. It might be this model is super good, I haven’t tried it, but to say the Chinese models are better is just not true. What I really love though is that I can run them (open models) on my own machine. The other day I categorised images locally using Qwen, what a time to be al…
For instance, a lot of people thought they were running "DeepSeek" when they were really running some random distillation on ollama.