Live data from Hacker News

DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

huggingface.co

161–170 of 485 posts

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#161
post #16

Well props to them for continuing to improve, winning on cost-effectiveness, and continuing to publicly share their improvements. Hard not to root for them as a force to prevent an AI corporate monopoly/duopoly.

[flagged]

Should I root for the democratic OpenAI, Google or Microsoft instead?

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#162
post #16

Well props to them for continuing to improve, winning on cost-effectiveness, and continuing to publicly share their improvements. Hard not to root for them as a force to prevent an AI corporate monopoly/duopoly.

>winning on cost-effectiveness

Nobody is winning in this area until these things run in full on single graphics cards. Which is sufficient compute to run even most of the complex tasks.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#163

At least, there is no doubt where he is from ! which version are you? 我是DeepSeek最新版本模型! 如果你想了解具体的版本号信息,我建议你: 查看官方文档 - DeepSeek官网和文档会有最准确的版本信息 关注官方公告 - 版本更新通常会在官方渠道公布 查看应用商店/网页版 - 使用界面通常会显示当前版本 我具备DeepSeek的所有最新功能特性,包括: 强大的对话和推理能力 128K上下文长度 文件上传处理(图像、文档等) 联网搜索功能(需手动开启) 完全免费使用 如果你需要知道精确的版本号用于技术对接或其他特定用途,最好直接查阅官方技术文档,那里会有最准确和详细的技术规格说明。 有什么其他问题我可以帮你解答吗?

It's so strange when it obviously hits a preprogrammed non-answer in these models, how can one ever trust them when there is a babysitter that interferes in an actual answer. I suppose that asking it what version it is isn't a valid question in it's training data so it's programmed to say check the documentation, but still definitely suspicious when it gives a non-answer.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#164

Earlier quoted context omitted.

on what hypothetical grounds would you be more meaningfully able to sue the american maker of a self-hosted statistical language model that you select your own runtime sampling parameters for after random subtle security vulnerabilities came out the other side when you asked it for very secure code? put another way, how do you propose to tell this subtle nefarious chinese sabotage you baselessly imply to be commonpla…

This paper may be of interest to you: https://arxiv.org/html/2504.15867v1

the mechanism of action for that attack appears to be reading from poisoned snippets on stackoverflow or a similar site, which to my mind is an excellent example of why it seems like it would be difficult to retroactively pin "insecure code came out of my model" on the evil communist base weights of the model in question

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#165

Earlier quoted context omitted.

I can’t think of a single company I’ve worked with as a consultant that I could convince to use DeepSeek because of its ties with China even if I explained that it was hosted on AWS and none of the information would go to China. Even when the technical people understood that, it would be too much of a political quagmire within their company when it became known to the higher ups. It just isn’t worth the political cap…

This is the real cause. At the enterprise level, trust outweighs cost. My company hires agencies and consultants who provide the same advice as our internal team; this is not to imply that our internal team is incorrect; rather, there is credibility that if something goes wrong, the decision consequences can be shifted, and there is a reason why companies continue to hire the same four consulting firms. It's trust, w…

So much worse for American companies. This only means that they will be uncompetitive with similar companies that use models with realistic costs.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#166
post #35
post #21

Earlier quoted context omitted.

For just the model itself: 4 x params at F32, 2 x params at F16/BF16, or 1 x params at F8, e.g. 685GB at F8. It will be smaller for quantizations, but I'm not sure how to estimate those. For a Mixture of Experts (MoE) model you only need to have the memory size of a given expert. There will be some swapping out as it figures out which expert to use, or to change expert, but once that expert is loaded it won't be swap…

I think your idea of MoE is incorrect. Despite the name they're not "expert" at anything in particular, used experts change more or less on each token -- so swapping them into VRAM is not viable, they just get executed on CPU (llama.cpp).

A common pattern is to offload (most of) the expert layers to the CPU. This combination is still quite fast even with slow system ram, though obviously inferior to a pure VRAM loading

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#167
post #9

It's awesome that stuff like this is open source, but even if you have a basement rig with 4 NVIDIA GeForce RTX 5090 graphic cards ($15-20k machine), can it even run with any reasonable context window that isn't like a crawling 10/tps? Frontier models are far exceeding even the most hardcore consumer hobbyist requirements. This is even further

FWIW it looks like OpenRouter's two providers for this model (one of whom being Deepseek itself) are only running the model around 28tps at the moment. https://openrouter.ai/deepseek/deepseek-v3.2 This only bolsters your point. Will be interesting to see if this changes as the model is adopted more widely.

[deleted]

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#168
post #55

Earlier quoted context omitted.

You can run at ~20 tokens/second on a 512GB Mac Studio M3 Ultra: https://youtu.be/ufXZI6aqOU8?si=YGowQ3cSzHDpgv4z&t=197 IIRC the 512GB mac studio is about $10k

and can be faster if you can get an MOE model of that

All modern models are MoE already, no?

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#169
post #129

Earlier quoted context omitted.

> It reminds me, in an encouraging way, of the way that German military planners regarded the Soviet Union in the lead-up to Operation Barbarossa. The Slavs are an obviously inferior race; their Bolshevism dooms them; we have the will to power; we will succeed Though, because Stalin had decimated the red army leadership (including most of the veteran officer who had Russian civil war experience) during the Moscow tri…

> Though, because Stalin had decimated the red army leadership (including most of the veteran officer who had Russian civil war experience) during the Moscow trials purges, the German almost succeeded. There were many counter revolutionaries among the leadership, even those conducting the purges. Stalin was like "ah fuck we're hella compromised." Many revolutions fail in this step and often end up facing a CIA backed…

> There were many counter revolutionaries among the leadership

Well, Stalin was, by far, the biggest counter-revolutionary in the Politburo.

> Stalin was like "ah fuck we're hella compromised."

There's no evidence that anything significant was compromised at that point, and clear evidence that Stalin was in fact medically paranoid.

> Many revolutions fail in this step and often end up facing a CIA backed coup. The USSR was under constant siege and attempted infiltration since inception.

Can we please not recycle 90-years old soviet propaganda? The Moscow trial being irrational self-harm was acknowledged by the USSR leadership as early as the fifties…

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#170
post #153

Earlier quoted context omitted.

on what hypothetical grounds would you be more meaningfully able to sue the american maker of a self-hosted statistical language model that you select your own runtime sampling parameters for after random subtle security vulnerabilities came out the other side when you asked it for very secure code? put another way, how do you propose to tell this subtle nefarious chinese sabotage you baselessly imply to be commonpla…

[flagged]

Competitor != adversary. It is US warmongering ideology that tries to equate these concepts.
Post reply on HN