Live data from Hacker News

DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

huggingface.co

431–440 of 485 posts

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#431

Earlier quoted context omitted.

I don't see what OpenAI's niche is supposed to be, other than role playing? Google seems like they'll be the AI utility company, and Anthropic seems like the go-to for the AI developer platform of the future.

Anthropic has RLed the shit out of their models to the extent that they give sub-par answers to general purpose questions. Google has great models but is institutionally incapable of building a cohesive product experience. They are literally shipping their org chart with Gemini (mediocre product), AI Overview (trash), AI Mode (outstanding but limited modality), Gemini for Google Workspace (steaming pile), Gemini on A…

Gemini is not mediocre, have you used it lately?

https://www.vellum.ai/llm-leaderboard

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#432

Earlier quoted context omitted.

Thanks for sharing that! The scales are a bit murky here, but if we look at the 'Coding' metric, we see that Kimi K2 outperforms Sonnet 4.5 - that's considered to be the price-perf darling I think even today? I haven't tried these models, but in general there have been lots of cases where a model performs much worse IRL than the benchmarks would sugges (certain Chinese models and GPT-OSS have been guilty of this in t…

Good question. There's 2 points to consider. • For both Kimi K2 and for Sonnet, there's a non-thinking and a thinking version. Sonnet 4.5 Thinking is better than Kimi K2 non-thinking, but the K2 Thinking model came out recently, and beats it on all comparable pure-coding benchmarks I know: OJ-Bench (Sonnet: 30.4% https://artificialanalysis.ai/models/capabilities/coding • The reason developers love Sonnet 4.5 for codi…

The table is confusing. It is not clear what is known and what is predicted (and how it is predicted). Why not measure the missing pieces instead of predicting—is it too expensive or is the tooling missing?

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#433

How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? What hurt open source in the past was its inability to keep up with the quality and feature depth of closed source competitors, but models seem to be reaching a performance plateau; the top open weight models are generally indistinguishable from the top private models. Infrastructure owners with acce…

>How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? According to Google (or someone at Google) no organization has moat on AI/LLM [1]. But that does not mean that it is not hugely profitable providing it as SaaS even you don't own the model or Model as a Service (MaaS). The extreme example is Amazon providing MongoDB API and services. Sure they have…

Hosting a SOTA AI model is something that can be separated well from the rest of your cloud deployments. So you can pretty much choose between lots of vendors and that means margins will probably not be that great.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#434
post #16

Well props to them for continuing to improve, winning on cost-effectiveness, and continuing to publicly share their improvements. Hard not to root for them as a force to prevent an AI corporate monopoly/duopoly.

How do they make their money

I suspect it is a state venture designed to undermine the American-led proprietary AI boom. I'm all for it, tbh, but as others have pointed out, if they successfully destroy the American ventures it's not like we can expect an altruistic endgame from them.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#435

Earlier quoted context omitted.

>How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? According to Google (or someone at Google) no organization has moat on AI/LLM [1]. But that does not mean that it is not hugely profitable providing it as SaaS even you don't own the model or Model as a Service (MaaS). The extreme example is Amazon providing MongoDB API and services. Sure they have…

That quote from Google is 2.5 years old.

undergrads at UC Berkeley are wearing vLLM t-shirts

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#436
post #120

Earlier quoted context omitted.

> I run it all the time, token generation is pretty good. I feel like because you didn't actually talk about prompt processing speed or token/s, you aren't really giving the whole picture here. What is the prompt processing tok/s and the generation tok/s actually like?

I addressed both points - I mentioned you can offload token prefill (the slow part, 9t/s) to DGX Spark. Token generation is at 6t/s which is acceptable.

6t/s will have you pulling your hair out with any deepseek model.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#437
post #324

Earlier quoted context omitted.

> I can’t think of a single major US company that is big internationally that is competing on price. All the clouds compete on price. Do you really think it is that differentiated? Google, Amazon and Microsoft all offer special deals to sign big companies up and globally too.

I worked inside AWS consulting department for 3 years (AWS ProServe) and now I work as a staff consultant for a 3rd AWS partner. I have been on enough sales calls, seen enough go to market training materials and flown out to customers sites to know how these things work. AWS has never tried to compete as the “low cost leader”. Marketing 101 says you never want to compete on price if you can avoid it. Microsoft doesn’…

> AWS has never tried to compete as the “low cost leader”. Marketing 101 says you never want to compete on price if you can avoid it.

Despite all that and whatever you say, the fact is you do compete. It doesn't have to be a race to the bottom.

So Cloudfront free tier and the latest discount bundles etc aren't to compete? People have also negotiated private pricing way below list price and a lot cheaper than competitors.

Similarly was the Dynamodb price cuts not due to competition?

I can give way more examples...

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#438
post #428

Earlier quoted context omitted.

Well the seemingly cheap comes with significantly degraded performance, particular for agentic use. Have you tried replacing Claude Code with some locally deployed model, say, on 4090 or 5090? I have. It is not usable.

Strictly speaking, you have not deployed any model on a 5090 because a 5090 card has never been produced. And without specifying your quantization level it's hard to know what you mean by "not usable" Anyway if you really wanted to try cheap distilled/quantized models locally you would be using used v100 Teslas and not 4 year old single chip gaming GPUs.

You can just buy a 5090 now for $3k. Have you confused it with something else?

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#439

How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? What hurt open source in the past was its inability to keep up with the quality and feature depth of closed source competitors, but models seem to be reaching a performance plateau; the top open weight models are generally indistinguishable from the top private models. Infrastructure owners with acce…

This is exactly why the CEO of Anthropic has been talking up "risks" from AI models and asking for legislation to regulate the industry.

He's talking about completely different type of risks and regulation. It's about the job displacement risks, security and misuse concerns, and ethical and societal impact.

https://www.youtube.com/watch?v=aAPpQC-3EyE

https://www.youtube.com/watch?v=RhOB3g0yZ5k

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#440
post #55
post #9

It's awesome that stuff like this is open source, but even if you have a basement rig with 4 NVIDIA GeForce RTX 5090 graphic cards ($15-20k machine), can it even run with any reasonable context window that isn't like a crawling 10/tps? Frontier models are far exceeding even the most hardcore consumer hobbyist requirements. This is even further

You can run at ~20 tokens/second on a 512GB Mac Studio M3 Ultra: https://youtu.be/ufXZI6aqOU8?si=YGowQ3cSzHDpgv4z&t=197 IIRC the 512GB mac studio is about $10k

~20 tokens/second is actually pretty good. I see he's using the q5 version of the model. I wonder how it scales with the larger contexts. And the same guy published the video today with the new 3.2 version: https://www.youtube.com/watch?v=b6RgBIROK5o
Post reply on HN