Live data from Hacker News

DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

huggingface.co

351–360 of 485 posts

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#351
post #51

Disclaimer: I did not test this yet. I don't want to make big generalizations. But one thing I noticed with chinese models, especially Kimi, is that it does very well on benchmarks, but fails on vibe testing. It feels a little bit over-fitting to the benchmark and less to the use cases. I hope it's not the same here.

My experience with deepseek and Kimi is quite the opposite: smarter than benchmarks would imply

Whereas the benchmark gains seem by new OpenAI, Grok and Claude models don't feel accompanied by vibe improvement

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#353
post #262

Earlier quoted context omitted.

How did we come to the place that the most transparent and open models are now coming out of China—freely sharing their research and source code—while all the American ones are fully locked down

Because the entire US economy is being propped up by AI hype.

The money would have gone somewhere. The "smart" money went to AI. Don't get fooled.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#354
post #5

I genuinely do not understand the evaluations of the US AI industry. The chinese models are so close and far cheaper

Two aspects to consider: 1. Chinese models typically focus on text. US and EU models also bear the cross of handling image, often voice and video. Supporting all those is additional training costs not spent on further reasoning, tying one hand in your back to be more generally useful. 2. The gap seems small, because so many benchmarks get saturated so fast. But towards the top, every 1% increase in benchmarks is sign…

> Chinese models typically focus on text

Not true at all. Qwen has a VLM (qwen2 vl instruct) which is the backbone of Bytedance’s TARS computer use model. Both Alibaba (Qwen) and Bytedance are Chinese.

Also DeepSeek got a ton of attention with their OCR paper a month ago which was an explicit example of using images rather than text.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#355

Earlier quoted context omitted.

Nobody is winning until cars are the size of a pack of cards. Which is big enough to transport even the largest cargo.

Lol its kinda suprising that the level of understanding around LLMs is so little. You already have agents, that can do a lot of "thinking", which is just generating guided context, then using that context to do tasks. You already have Vector Databases that are used as context stores with information retrieval. Fundamentally, you can have the same exact performance on a lot of task whether all the information exists i…

You should start a company and try your strategy. I hope it works! (Though I am doubtful.)

In any case, models are useful, even when they don't hit these efficiency targets you are projecting. Just like cars are useful, even when they are bigger than a pack of cards.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#357

Earlier quoted context omitted.

Why does that matter? They wont be making at home graphics cards anymore. Why would you do that when you can be pre-sold $40k servers for years into the future

Because Moore's law marches on. We're around 35-40 orders of magnitude from computers now to computronium. We'll need 10-15 years before handheld devices can run a couple terabytes of ram, 64-128 terabytes of storage, and 80+ TFLOPS. That's enough to run any current state of the art AI at around 50 tokens per second, but in 10 years, we're probably going to have seen lots of improvements, so I'd guess conservatively…

> What if you had the equivalent of a frontier lab in your pocket? What's that do to the economy?

Well, these days people have the equivalent of a frontier lab from perhaps 40 years ago in their pocket. We can see what that has done to the economy, and try to extrapolate.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#358
post #16

Well props to them for continuing to improve, winning on cost-effectiveness, and continuing to publicly share their improvements. Hard not to root for them as a force to prevent an AI corporate monopoly/duopoly.

How could we judge if anyone is "winning" on cost-effectiveness, when we don't know what everyones profits/losses are?

Apart from measuring prices from venture-backed providers which might or might not correlate with cost-effectiveness, I think the measures of intelligence per watt and intelligence per joule from https://arxiv.org/abs/2511.07885 is very interesting.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#359

Earlier quoted context omitted.

I appreciate your rabid optimism, but considering that Moores Law has ceased to be true for multiple years now I am not sure a handwave about being able to scale to infinity is a reasonable way to look at things. Plenty of things have slowed down in progress in our current age, for example airplanes.

Someone always crawls out of the woodwork to repeat this supposed "fact" which hasn't been true for the entire half-century it's been repeated. Jim Keller (designer of most of the great CPUs of the last couple decades) gave a convincing presentation several years ago about just how not-true it is: https://www.youtube.com/watch?v=oIG9ztQw2Gc Everything he says in it still applies today. Intel struggled for a decade, a…

During the 1990s (and for some years before and after) we got 'Dennard scaling'. The frequency of processors tended to increase exponentially, too, and featured prominently in advertising and branding.

I suspect many people conflated Dennard scaling with Moore's law and the demise of Dennard scaling is what contributes to the popular imagination that Moore's law is dead: frequencies of processors have essentially stagnated.

See https://en.wikipedia.org/wiki/Dennard_scaling

Post reply on HN