Live data from Hacker News

DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

huggingface.co

61–70 of 485 posts

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#61
post #51

Disclaimer: I did not test this yet. I don't want to make big generalizations. But one thing I noticed with chinese models, especially Kimi, is that it does very well on benchmarks, but fails on vibe testing. It feels a little bit over-fitting to the benchmark and less to the use cases. I hope it's not the same here.

This used to happen with bench marks on phones, manufacturers would tweak android so benchmarks ran faster. I guess that’s kinda how it is for any system that’s trained to do well on benchmarks, it does well but rubbish at everything else.

yes, they turned off all energy economy measures when benchmarking software activity was detected, which completely broke the point of the benchmarks because your phone is useless if it's very fast but the battery lasts one hour

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#62
I hate that their model ids don't change as they change the underlying model. I'm not sure how you can build on that.

  % curl https://api.deepseek.com/models \          
    -H "Authorization: Bearer ${DEEPSEEK_API_KEY}"  
  {"object":"list","data":[{"id":"deepseek-chat","object":"model","owned_by":"deepseek"},{"id":"deepseek-reasoner","object":"model","owned_by":"deepseek"}]}

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#63
post #51

Disclaimer: I did not test this yet. I don't want to make big generalizations. But one thing I noticed with chinese models, especially Kimi, is that it does very well on benchmarks, but fails on vibe testing. It feels a little bit over-fitting to the benchmark and less to the use cases. I hope it's not the same here.

What is "Vibe testing"?

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#64
post #10
post #5

I genuinely do not understand the evaluations of the US AI industry. The chinese models are so close and far cheaper

It's all about the hardware and infrastructure. If you check OpenRouter, no provider offers a SOTA chinese model matching the speed of Claude, GPT or Gemini. The chinese models may benchmark close on paper, but real-world deployment is different. So you either buy your own hardware in order to run a chinese model at 150-200tps or give up an use one of the Big 3. The US labs aren't just selling models, they're selling…

Gemini 3 = ~70tps https://openrouter.ai/google/gemini-3-pro-preview

Opus 4.5 = ~60-80tps https://openrouter.ai/anthropic/claude-opus-4.5

Kimi-k2-think = ~60-180tps https://openrouter.ai/moonshotai/kimi-k2-thinking

Deepseek-v3.2 = ~30-110tps (only 2 providers rn) https://openrouter.ai/deepseek/deepseek-v3.2

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#65
post #29
post #20

Earlier quoted context omitted.

> Valuation is not based on what they have done but what they might do Exactly what I’m thinking. Chinese models catching rapidly. Soon to be on-par with the big dogs.

Even if they do continue to lag behind they are a good bet against monopolisation by proprietary vendors.

They would if corporations were allowed to run these models. I fully expect the US government to prohibit corporations from doing anything useful with Chinese models (full censorship). It's the same game they use with chips.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#67
post #16

Well props to them for continuing to improve, winning on cost-effectiveness, and continuing to publicly share their improvements. Hard not to root for them as a force to prevent an AI corporate monopoly/duopoly.

I suspect they will keep doing this until they have a substantially better model than the competition. Sharing methods to look good & allow the field to help you keep up with the big guys is easy. I'll be impressed if they keep publishing even when they do beat the big guys soundly.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#68
post #51

Disclaimer: I did not test this yet. I don't want to make big generalizations. But one thing I noticed with chinese models, especially Kimi, is that it does very well on benchmarks, but fails on vibe testing. It feels a little bit over-fitting to the benchmark and less to the use cases. I hope it's not the same here.

What is "Vibe testing"?

I would assume that it is testing how well and appropriately the LLM responds to prompts.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#69
post #12

Earlier quoted context omitted.

Valuation is not based on what they have done but what they might do, I agree tho it's investment made with very little insight into Chinese research. I guess it's counting on deepseek being banned and all computers in America refusing to run open software by the year 2030 /snark

> I guess it's counting on deepseek being banned And the people making the bets are in a position to make sure the banning happens. The US government system being what it is. Not that our leaders need any incentive to ban Chinese tech in this space. Just pointing out that it's not necessarily a "bet". "Bet" imply you don't know the outcome and you have no influence over the outcome. Even "investment" implies you don'…

Exactly. "Business investment" these days means that the people involved will have at least some amount of power to determine the winning results.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#70

Earlier quoted context omitted.

Two aspects to consider: 1. Chinese models typically focus on text. US and EU models also bear the cross of handling image, often voice and video. Supporting all those is additional training costs not spent on further reasoning, tying one hand in your back to be more generally useful. 2. The gap seems small, because so many benchmarks get saturated so fast. But towards the top, every 1% increase in benchmarks is sign…

Nothing you said helps with the issue of valuation. Yes, the US models may be better by a few percentage points, but how can they justify being so costly, both operationally as well as in investment costs? Over the long run, this is a business and you don't make money being the first, you have to be more profitable overall.

I think the investment race here is an "all-pay auction"*. Lots of investors have looked at the ultimate prize — basically winning something larger than the entire present world economy forever — and think "yes".

But even assuming that we're on the right path for that (which we may not be) and assuming that nothing intervenes to stop it (which it might), there may be only one winner, and that winner may not have even entered the game yet.

* https://en.wikipedia.org/wiki/All-pay_auction

Post reply on HN