Earlier quoted context omitted.
Sparse Attention, it's the highlight of this model as per the paper
How did we come to the place that the most transparent and open models are now coming out of China—freely sharing their research and source code—while all the American ones are fully locked down
DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
301–310 of 485 posts
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#302Earlier quoted context omitted.
This is the rare earth minerals dumping all over again. Devalue to such a price as to make the market participants quit, so they can later have a strategic stranglehold on the supply. This is using open source in a bit of different spirit than the hacker ethos, and I am not sure how I feel about it. It is a kind of cheat on the fair market but at the same time it is also costly to China and its capital costs may beco…
Ah, so exactly like Uber, Netflix, Microsoft, Amazon, Facebook and so on have done to the rest of the world over the last few decades then? Where do you think they learnt this trick? Years lurking on HN and this post's comment section wins #1 on the American Hypocrisy chart. Unbelievable that even in the current US people can't recognize when they're looking in the mirror. But I guess you're disincentivized to do so…
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#303How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? What hurt open source in the past was its inability to keep up with the quality and feature depth of closed source competitors, but models seem to be reaching a performance plateau; the top open weight models are generally indistinguishable from the top private models. Infrastructure owners with acce…
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#3043.2-Exp came out in September: this is 3.2, along with a special checkpoint (DeepSeek-V3.2-Speciale) for deep reasoning that they're claiming surpasses GPT-5 and matches Gemini 3.0 https://x.com/deepseek_ai/status/1995452641430651132
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#305Earlier quoted context omitted.
If you're trying to build AI based applications you can and should compare the costs between vendor based solutions and hosting open models with your own hardware. On the hardware side you can run some benchmarks on the hardware (or use other people's benchmarks) and get an idea of the tokens/second you can get from the machine. Normalize this for your usage pattern (and do your best to implement batch processing whe…
Well the seemingly cheap comes with significantly degraded performance, particular for agentic use. Have you tried replacing Claude Code with some locally deployed model, say, on 4090 or 5090? I have. It is not usable.
Things get a lot more easier at lower quantisation, higher parameter space, and there's a lot of people's whose jobs for AI are "Extract sentiment from text" or "bin into one of these 5 categories" where that's probably fine.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#306How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? What hurt open source in the past was its inability to keep up with the quality and feature depth of closed source competitors, but models seem to be reaching a performance plateau; the top open weight models are generally indistinguishable from the top private models. Infrastructure owners with acce…
People and companies trust OpenAI and Anthropic, rightly or wrongly, with hosting the models and keeping their company data secure. Don't underestimate the value of a scapegoat to point a finger at when things go wrong.
Which is an interesting point in favour of the human employee, as you can only consolidate scape goats so far up the chain before saying "It was AIs fault" just looks like negligence.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#307Earlier quoted context omitted.
I can’t think of a single company I’ve worked with as a consultant that I could convince to use DeepSeek because of its ties with China even if I explained that it was hosted on AWS and none of the information would go to China. Even when the technical people understood that, it would be too much of a political quagmire within their company when it became known to the higher ups. It just isn’t worth the political cap…
AirBnB is all in on DeepSeek and Qwen. https://sg.finance.yahoo.com/news/airbnb-picks-alibabas-qwen...
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#308Earlier quoted context omitted.
How could we judge if anyone is "winning" on cost-effectiveness, when we don't know what everyones profits/losses are?
Well consumers care about the cost to them, and those we know. And deepseek is destroying everything in that department.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#309Earlier quoted context omitted.
The performance bottleneck for space based computers is heat dissipation. Earth based computers benefit from the existence of an atmosphere to pull cold air in from and send hot air out to. A space data center would need to entirely rely on city sized heat sink fins.
For radiative cooling using aluminum, per 1000 watts at 300 kelvin: ~2.4m^2 area, ~4.8 liters volume, ~13kg weight. So a Starship (150k kg, re-usable) could carry about a megawatt of radiators per launch to LEO. And aluminum is abundant in the lunar crust.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#310How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? What hurt open source in the past was its inability to keep up with the quality and feature depth of closed source competitors, but models seem to be reaching a performance plateau; the top open weight models are generally indistinguishable from the top private models. Infrastructure owners with acce…
hopefully they won't
and their titanic off-balance sheet investments will bankrupt them as they won't be able to produce any revenue