Live data from Hacker News

DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

huggingface.co

441–450 of 485 posts

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#441

Can someone kind please ELI5 this paper?

They've developed a sparse attention mechanism (which they document and release source code for) to increase model efficiency with long context, as needed for fast & cost-effective extensive RL training for reasoning and agentic use

They've built a "stable & scalable" RL protocol - more capable RL training infrastructure

They've built a pipeline/process to generate synthetic data for reasoning and agentic training

These all combine to build an efficient model with extensive RL post-training for reasoning and agentic use, although they note work is still needed on both the base model (more knowledge) and post-training to match frontier performance.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#442

Earlier quoted context omitted.

Lol its kinda suprising that the level of understanding around LLMs is so little. You already have agents, that can do a lot of "thinking", which is just generating guided context, then using that context to do tasks. You already have Vector Databases that are used as context stores with information retrieval. Fundamentally, you can have the same exact performance on a lot of task whether all the information exists i…

https://en.wikipedia.org/wiki/Dunning%E2%80%93Kruger_effect

Ok then point out where I made a mistake.

Nothing shows lack of understanding of the subject matter more than referencing the Dunning Kruger effect in a conversation.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#443
post #416

How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? What hurt open source in the past was its inability to keep up with the quality and feature depth of closed source competitors, but models seem to be reaching a performance plateau; the top open weight models are generally indistinguishable from the top private models. Infrastructure owners with acce…

Yes but how do you find the best open model? You check google.

Kagi

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#444

Earlier quoted context omitted.

Because Moore's law marches on. We're around 35-40 orders of magnitude from computers now to computronium. We'll need 10-15 years before handheld devices can run a couple terabytes of ram, 64-128 terabytes of storage, and 80+ TFLOPS. That's enough to run any current state of the art AI at around 50 tokens per second, but in 10 years, we're probably going to have seen lots of improvements, so I'd guess conservatively…

I appreciate your rabid optimism, but considering that Moores Law has ceased to be true for multiple years now I am not sure a handwave about being able to scale to infinity is a reasonable way to look at things. Plenty of things have slowed down in progress in our current age, for example airplanes.

The Law of Accelerating Returns is a better formulation, not tied to any particular substrate, it's just not as widely known.

https://imgur.com/a/UOUGYzZ - had chatgpt whip up an updated chart.

LoAR shows remarkably steady improvement. It's not about space or power efficiency, just ops per $1000, so transistor counts served as a very good proxy for a long time.

There's been sufficiently predictable progress that 80-100 TFLOPS in your pocket by 3035 is probably a solid bet, especially if a fully generative AI OS and platform catches on as a product. The LoAR frontier for compute in 2035 is going to be more advanced than the limits of prosumer/flagship handheld products like phones, so theres a bit of lag and variability.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#445

Earlier quoted context omitted.

> There are plenty of 3rd party and big cloud options to run these models by the hour or token. Which ones? I wanted to try a large base model for automated literature (fine-tuned models are a lot worse at it) but I couldn't find a provider which makes this easy.

Fireworks supports this model serverless for $1.20 per million tokens. https://fireworks.ai/models/fireworks/deepseek-v3p2

That's the final, fine-tuned model. The base model (pretraining only, no instruction SFT, RLHF, RLVR etc) is this one: https://huggingface.co/deepseek-ai/DeepSeek-V3.2-Exp-Base It's apparently not offered at any inference provider, nor are older DeepSeek base models.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#446

How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? What hurt open source in the past was its inability to keep up with the quality and feature depth of closed source competitors, but models seem to be reaching a performance plateau; the top open weight models are generally indistinguishable from the top private models. Infrastructure owners with acce…

> How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models?

So a couple of things. There are going to be a handful of companies in the world with the infrastructure footprint and engineering org capable of running LLMs efficiently and at scale. You are never going to be able to run open models in your own infra in a way that is cost competitive with using their API.

Competition _between_ the largest AI companies _will_ drive API prices to essentially 0 profit margin, but none of those companies will care because they aren't primarily going to make money by selling the LLM API -- your usage of their API just subsidizes their infrastructure costs, and they'll use that infra to build products like chat gpt and claude, etc. Those products are their moat and will be where 90% of their profit comes from.

I am not sure why everyone is so obsessed with "moats" anyway. Why does gmail have so many users? Anybody can build an email app. For the same reason that people stick with gmail, people are going to stick with chatgpt. It's being integrated into every aspect of their lives. The switching costs for people are going to be immense.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#447

Earlier quoted context omitted.

Eh, I don't see the risk, no pun intended. It's not collimated, and it's not going to be in focus anywhere but on-target. It's also probably in the long-wave range >>1000 nm that's not focused by the eye. At the end of the day it's no different from any other source of spot heating. I get more nervous around some of the LED flashlights you can buy these days. I want one. Hot air blows.

It's 45w of lasing power. I have a scar on my hand that's 15 years old from running one of those at 10% power and getting a reflection from a bare metal sheet. This will absolutely scar, if not char, your cornea faster than you can blink.

That's (again) less energy than a flashlight puts out these days, so the beam had to be tightly focused in your case. That isn't how these things work.

There is nothing special about "lasing power." It amounts to a 45-watt light bulb, nothing more and nothing less.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#448
The AI market is hard to predict due to the constant development of new algorithms that could emerge unexpectedly. Refer to this summary of Ilya's opinions for insights into the necessity of these new algorithms: https://youtu.be/DcrXHTOxi3I

DeepSeek is a valuable product, but its open-source nature makes it difficult to displace larger competitors. Any advancements can be quickly adopted, and in fact, it may inadvertently strengthen these companies by highlighting weaknesses in their current strategies.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#449
post #399

Earlier quoted context omitted.

Deepseek and Kimi both have great agentic performance When used with crush/opencode they are close to Claude performance. Nothing that runs on a 4090 would compete but Deepseek on openrouter is still 25x cheaper than claude

> Deepseek on openrouter is still 25x cheaper than claude Is it? Or only when you don’t factor in Claude cached context? I’ve consistently found it pointless to use open models because the price of the good ones is so close to cached context on Claude that I don’t need them.

Yes, if you try using Kilo Code/Cline via Openrouter the cost will be much cheaper using Deepseek/Kimi vs Claude Sonnet 4.5.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#450

Why are there so few 32,64,128,256,512 GB models which could run on current consumer hardware? And why is the maximum RAM on Mac studio M4 128 GB??

the only real benefit is privacy which 99.9% of people dont get about. Almost all serving metrics (cost, throughput, ttft) are better with large gpu clusters. Latency is usually hidden by prefill cost.

and sovereignty. I can go into the woods with a fuzzy approximation of all internet text in my backpack
Post reply on HN