It's awesome that stuff like this is open source, but even if you have a basement rig with 4 NVIDIA GeForce RTX 5090 graphic cards ($15-20k machine), can it even run with any reasonable context window that isn't like a crawling 10/tps? Frontier models are far exceeding even the most hardcore consumer hobbyist requirements. This is even further
DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
11–20 of 485 posts
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#12I genuinely do not understand the evaluations of the US AI industry. The chinese models are so close and far cheaper
Valuation is not based on what they have done but what they might do, I agree tho it's investment made with very little insight into Chinese research. I guess it's counting on deepseek being banned and all computers in America refusing to run open software by the year 2030 /snark
And the people making the bets are in a position to make sure the banning happens. The US government system being what it is.
Not that our leaders need any incentive to ban Chinese tech in this space. Just pointing out that it's not necessarily a "bet".
"Bet" imply you don't know the outcome and you have no influence over the outcome. Even "investment" implies you don't know the outcome. I'm not sure that's the case with these people?
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#13Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#14It's awesome that stuff like this is open source, but even if you have a basement rig with 4 NVIDIA GeForce RTX 5090 graphic cards ($15-20k machine), can it even run with any reasonable context window that isn't like a crawling 10/tps? Frontier models are far exceeding even the most hardcore consumer hobbyist requirements. This is even further
People with basement rigs generally aren't the target audience for these gigantic models. You'd get much better results out of an MoE model like Qwen3's A3B/A22B weights, if you're running a homelab setup.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#15I genuinely do not understand the evaluations of the US AI industry. The chinese models are so close and far cheaper
It's all about the hardware and infrastructure. If you check OpenRouter, no provider offers a SOTA chinese model matching the speed of Claude, GPT or Gemini. The chinese models may benchmark close on paper, but real-world deployment is different. So you either buy your own hardware in order to run a chinese model at 150-200tps or give up an use one of the Big 3. The US labs aren't just selling models, they're selling…
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#16Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#17Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#18I genuinely do not understand the evaluations of the US AI industry. The chinese models are so close and far cheaper
It reminds me, in an encouraging way, of the way that German military planners regarded the Soviet Union in the lead-up to Operation Barbarossa. The Slavs are an obviously inferior race; their Bolshevism dooms them; we have the will to power; we will succeed. Even now, when you ask questions like what you ask of that era, the answers you get are genuinely not better than "yes, this should have been obvious at the time if you were not completely blinded by ethnic and especially ideological prejudice."
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#19I genuinely do not understand the evaluations of the US AI industry. The chinese models are so close and far cheaper
There is a great deal of orientalism --- it is genuinely unthinkable to a lot of American tech dullards that the Chinese could be better at anything requiring what they think of as "intelligence." Aren't they Communist? Backward? Don't they eat weird stuff at wet markets? It reminds me, in an encouraging way, of the way that German military planners regarded the Soviet Union in the lead-up to Operation Barbarossa. Th…
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#20I genuinely do not understand the evaluations of the US AI industry. The chinese models are so close and far cheaper
Valuation is not based on what they have done but what they might do, I agree tho it's investment made with very little insight into Chinese research. I guess it's counting on deepseek being banned and all computers in America refusing to run open software by the year 2030 /snark
Exactly what I’m thinking. Chinese models catching rapidly. Soon to be on-par with the big dogs.