Live data from Hacker News

DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

huggingface.co

301–310 of 485 posts

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#301
post #262

Earlier quoted context omitted.

Sparse Attention, it's the highlight of this model as per the paper

How did we come to the place that the most transparent and open models are now coming out of China—freely sharing their research and source code—while all the American ones are fully locked down

Because the whole framing of US vs China as open vs closed was never correct to begin with.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#302
post #261

Earlier quoted context omitted.

This is the rare earth minerals dumping all over again. Devalue to such a price as to make the market participants quit, so they can later have a strategic stranglehold on the supply. This is using open source in a bit of different spirit than the hacker ethos, and I am not sure how I feel about it. It is a kind of cheat on the fair market but at the same time it is also costly to China and its capital costs may beco…

Ah, so exactly like Uber, Netflix, Microsoft, Amazon, Facebook and so on have done to the rest of the world over the last few decades then? Where do you think they learnt this trick? Years lurking on HN and this post's comment section wins #1 on the American Hypocrisy chart. Unbelievable that even in the current US people can't recognize when they're looking in the mirror. But I guess you're disincentivized to do so…

Except domestic alternatives to the tech companies you listed were not driven out by them, they still exist today with substantial market share. American tech dominance elsewhere has more to do a lack of competition, and when competition does exist they're more often than not held at a disadvantage by domestic governments. So your counter narrative is false here.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#303

How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? What hurt open source in the past was its inability to keep up with the quality and feature depth of closed source competitors, but models seem to be reaching a performance plateau; the top open weight models are generally indistinguishable from the top private models. Infrastructure owners with acce…

I don't see what OpenAI's niche is supposed to be, other than role playing? Google seems like they'll be the AI utility company, and Anthropic seems like the go-to for the AI developer platform of the future.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#304

3.2-Exp came out in September: this is 3.2, along with a special checkpoint (DeepSeek-V3.2-Speciale) for deep reasoning that they're claiming surpasses GPT-5 and matches Gemini 3.0 https://x.com/deepseek_ai/status/1995452641430651132

The assumption here is that 3.2 (without suffix) is an evolution of 3.2-Exp rather than being the same model, but they don't seem to be explicitly stating anywhere whether they're actually different or that they just made the same model GA.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#305
post #181

Earlier quoted context omitted.

If you're trying to build AI based applications you can and should compare the costs between vendor based solutions and hosting open models with your own hardware. On the hardware side you can run some benchmarks on the hardware (or use other people's benchmarks) and get an idea of the tokens/second you can get from the machine. Normalize this for your usage pattern (and do your best to implement batch processing whe…

Well the seemingly cheap comes with significantly degraded performance, particular for agentic use. Have you tried replacing Claude Code with some locally deployed model, say, on 4090 or 5090? I have. It is not usable.

Well, those are also extremely limited vram areas that wouldn't be able to run anything in the ~70b parameter space. (Can you run 30b even?)

Things get a lot more easier at lower quantisation, higher parameter space, and there's a lot of people's whose jobs for AI are "Extract sentiment from text" or "bin into one of these 5 categories" where that's probably fine.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#306

How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? What hurt open source in the past was its inability to keep up with the quality and feature depth of closed source competitors, but models seem to be reaching a performance plateau; the top open weight models are generally indistinguishable from the top private models. Infrastructure owners with acce…

People and companies trust OpenAI and Anthropic, rightly or wrongly, with hosting the models and keeping their company data secure. Don't underestimate the value of a scapegoat to point a finger at when things go wrong.

> Don't underestimate the value of a scapegoat to point a finger at when things go wrong.

Which is an interesting point in favour of the human employee, as you can only consolidate scape goats so far up the chain before saying "It was AIs fault" just looks like negligence.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#307

Earlier quoted context omitted.

I can’t think of a single company I’ve worked with as a consultant that I could convince to use DeepSeek because of its ties with China even if I explained that it was hosted on AWS and none of the information would go to China. Even when the technical people understood that, it would be too much of a political quagmire within their company when it became known to the higher ups. It just isn’t worth the political cap…

AirBnB is all in on DeepSeek and Qwen. https://sg.finance.yahoo.com/news/airbnb-picks-alibabas-qwen...

TIL: That Chinese models are considered better at multiple languages than non Chinese models.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#308

Earlier quoted context omitted.

How could we judge if anyone is "winning" on cost-effectiveness, when we don't know what everyones profits/losses are?

Well consumers care about the cost to them, and those we know. And deepseek is destroying everything in that department.

Yes. Though we don't know for sure whether that's because they actually have lower costs, or whether it's just the Chinese taxpayer being forced to serve us a treat.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#309

Earlier quoted context omitted.

The performance bottleneck for space based computers is heat dissipation. Earth based computers benefit from the existence of an atmosphere to pull cold air in from and send hot air out to. A space data center would need to entirely rely on city sized heat sink fins.

For radiative cooling using aluminum, per 1000 watts at 300 kelvin: ~2.4m^2 area, ~4.8 liters volume, ~13kg weight. So a Starship (150k kg, re-usable) could carry about a megawatt of radiators per launch to LEO. And aluminum is abundant in the lunar crust.

We are jumping pretty far ahead for a planet that can barely put two humans up there, but it is a great deal of my scifi dreams in one technology tree so I'll happily watch them try.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#310

How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? What hurt open source in the past was its inability to keep up with the quality and feature depth of closed source competitors, but models seem to be reaching a performance plateau; the top open weight models are generally indistinguishable from the top private models. Infrastructure owners with acce…

> How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models?

hopefully they won't

and their titanic off-balance sheet investments will bankrupt them as they won't be able to produce any revenue

Post reply on HN