Live data from Hacker News

DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

huggingface.co

81–90 of 485 posts

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#82
post #5

I genuinely do not understand the evaluations of the US AI industry. The chinese models are so close and far cheaper

They're not that close (on things like LMArena) and being cheaper is pretty meaningless when we are not yet at the point where LLMs are good enough for autonomy.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#84

I hate that their model ids don't change as they change the underlying model. I'm not sure how you can build on that. % curl https://api.deepseek.com/models \ -H "Authorization: Bearer ${DEEPSEEK_API_KEY}" {"object":"list","data":[{"id":"deepseek-chat","object":"model","owned_by":"deepseek"},{"id":"deepseek-reasoner","object":"model","owned_by":"deepseek"}]}

Oh hey, quality improvement without doing anything!

(unless/until a new version gets worse for your use case)

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#85
post #16

Well props to them for continuing to improve, winning on cost-effectiveness, and continuing to publicly share their improvements. Hard not to root for them as a force to prevent an AI corporate monopoly/duopoly.

As much I agree with your sentiment, but I doubt the intention is singular.

I don't care if this kills Google and OpenAI.

I hope it does, though I'm doubtful because distribution is important. You can't beat "ChatGPT" as a brand in laypeople's minds (unless perhaps you give them a massive "Temu: Shop Like A Billionaire" commercial campaign).

Closed source AI is almost by design morphing into an industrial, infrastructure-heavy rocket science that commoners can't keep up with. The companies pushing it are building an industry we can't participate or share in. They're cordoning off areas of tech and staking ground for themselves. It's placing a steep fence around tech.

I hope every such closed source AI effort is met with equivalent open source and that the investments made into closed AI go to zero.

The most likely outcome is that Google, OpenAI, and Anthropic win and every other "lab"-shaped company dies an expensive death. RunwayML spent hundreds of millions and they're barely noticeable now.

These open source models hasten the deaths of the second tier also-ran companies. As much as I hope for dents in the big three, I'm doubtful.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#86
post #5

I genuinely do not understand the evaluations of the US AI industry. The chinese models are so close and far cheaper

I would expect one of the motivations for making these LLM model weights open is to undermine the valuation of other players in the industry. Open models like this must diminish the value prop of the frontier focused companies if other companies can compete with similar results at competitive prices.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#87
post #16

Well props to them for continuing to improve, winning on cost-effectiveness, and continuing to publicly share their improvements. Hard not to root for them as a force to prevent an AI corporate monopoly/duopoly.

As much I agree with your sentiment, but I doubt the intention is singular.

The bar is incredibly low considering what OpenAI has done as a "not for profit"

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#88

To push back on naivety I'm sensing here I think it's a little silly to see Chinese Communist Party backed enterprise as somehow magnanimous and without ulterior, very harmful motive.

Oh they need control of models to be able to censor and ensure whatever happens inside the country with AI stays under their control. But the open-source part? Idk I think they do it to mess with the US investment and for the typical open source reasons of companies: community, marketing, etc. But tbh especially the messing with the US, as a european with no serious competitor, I can get behind.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#89
post #9

It's awesome that stuff like this is open source, but even if you have a basement rig with 4 NVIDIA GeForce RTX 5090 graphic cards ($15-20k machine), can it even run with any reasonable context window that isn't like a crawling 10/tps? Frontier models are far exceeding even the most hardcore consumer hobbyist requirements. This is even further

There are plenty of 3rd party and big cloud options to run these models by the hour or token. Big models really only work in that context, and that’s ok. Or you can get yourself an H100 rack and go nuts, but there is little downside to using a cloud provider on a per-token basis.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#90

To push back on naivety I'm sensing here I think it's a little silly to see Chinese Communist Party backed enterprise as somehow magnanimous and without ulterior, very harmful motive.

Oh they need control of models to be able to censor and ensure whatever happens inside the country with AI stays under their control. But the open-source part? Idk I think they do it to mess with the US investment and for the typical open source reasons of companies: community, marketing, etc. But tbh especially the messing with the US, as a european with no serious competitor, I can get behind.

They're pouring money to disrupt American AI markets and efforts. They do this in countless other fields. It's a model of massive state funding -> give it away for cut-rate -> dominate the market -> reap the rewards.

It's a very transparent, consistent strategy.

AI is a little different because it has geopolitical implications.

Post reply on HN