Earlier quoted context omitted.
I don't see what OpenAI's niche is supposed to be, other than role playing? Google seems like they'll be the AI utility company, and Anthropic seems like the go-to for the AI developer platform of the future.
Anthropic has RLed the shit out of their models to the extent that they give sub-par answers to general purpose questions. Google has great models but is institutionally incapable of building a cohesive product experience. They are literally shipping their org chart with Gemini (mediocre product), AI Overview (trash), AI Mode (outstanding but limited modality), Gemini for Google Workspace (steaming pile), Gemini on A…
DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
431–440 of 485 posts
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#432Earlier quoted context omitted.
Thanks for sharing that! The scales are a bit murky here, but if we look at the 'Coding' metric, we see that Kimi K2 outperforms Sonnet 4.5 - that's considered to be the price-perf darling I think even today? I haven't tried these models, but in general there have been lots of cases where a model performs much worse IRL than the benchmarks would sugges (certain Chinese models and GPT-OSS have been guilty of this in t…
Good question. There's 2 points to consider. • For both Kimi K2 and for Sonnet, there's a non-thinking and a thinking version. Sonnet 4.5 Thinking is better than Kimi K2 non-thinking, but the K2 Thinking model came out recently, and beats it on all comparable pure-coding benchmarks I know: OJ-Bench (Sonnet: 30.4% https://artificialanalysis.ai/models/capabilities/coding • The reason developers love Sonnet 4.5 for codi…
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#433How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? What hurt open source in the past was its inability to keep up with the quality and feature depth of closed source competitors, but models seem to be reaching a performance plateau; the top open weight models are generally indistinguishable from the top private models. Infrastructure owners with acce…
>How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? According to Google (or someone at Google) no organization has moat on AI/LLM [1]. But that does not mean that it is not hugely profitable providing it as SaaS even you don't own the model or Model as a Service (MaaS). The extreme example is Amazon providing MongoDB API and services. Sure they have…
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#434Well props to them for continuing to improve, winning on cost-effectiveness, and continuing to publicly share their improvements. Hard not to root for them as a force to prevent an AI corporate monopoly/duopoly.
How do they make their money
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#435Earlier quoted context omitted.
>How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? According to Google (or someone at Google) no organization has moat on AI/LLM [1]. But that does not mean that it is not hugely profitable providing it as SaaS even you don't own the model or Model as a Service (MaaS). The extreme example is Amazon providing MongoDB API and services. Sure they have…
That quote from Google is 2.5 years old.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#436Earlier quoted context omitted.
> I run it all the time, token generation is pretty good. I feel like because you didn't actually talk about prompt processing speed or token/s, you aren't really giving the whole picture here. What is the prompt processing tok/s and the generation tok/s actually like?
I addressed both points - I mentioned you can offload token prefill (the slow part, 9t/s) to DGX Spark. Token generation is at 6t/s which is acceptable.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#437Earlier quoted context omitted.
> I can’t think of a single major US company that is big internationally that is competing on price. All the clouds compete on price. Do you really think it is that differentiated? Google, Amazon and Microsoft all offer special deals to sign big companies up and globally too.
I worked inside AWS consulting department for 3 years (AWS ProServe) and now I work as a staff consultant for a 3rd AWS partner. I have been on enough sales calls, seen enough go to market training materials and flown out to customers sites to know how these things work. AWS has never tried to compete as the “low cost leader”. Marketing 101 says you never want to compete on price if you can avoid it. Microsoft doesn’…
Despite all that and whatever you say, the fact is you do compete. It doesn't have to be a race to the bottom.
So Cloudfront free tier and the latest discount bundles etc aren't to compete? People have also negotiated private pricing way below list price and a lot cheaper than competitors.
Similarly was the Dynamodb price cuts not due to competition?
I can give way more examples...
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#438Earlier quoted context omitted.
Well the seemingly cheap comes with significantly degraded performance, particular for agentic use. Have you tried replacing Claude Code with some locally deployed model, say, on 4090 or 5090? I have. It is not usable.
Strictly speaking, you have not deployed any model on a 5090 because a 5090 card has never been produced. And without specifying your quantization level it's hard to know what you mean by "not usable" Anyway if you really wanted to try cheap distilled/quantized models locally you would be using used v100 Teslas and not 4 year old single chip gaming GPUs.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#439How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? What hurt open source in the past was its inability to keep up with the quality and feature depth of closed source competitors, but models seem to be reaching a performance plateau; the top open weight models are generally indistinguishable from the top private models. Infrastructure owners with acce…
This is exactly why the CEO of Anthropic has been talking up "risks" from AI models and asking for legislation to regulate the industry.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#440It's awesome that stuff like this is open source, but even if you have a basement rig with 4 NVIDIA GeForce RTX 5090 graphic cards ($15-20k machine), can it even run with any reasonable context window that isn't like a crawling 10/tps? Frontier models are far exceeding even the most hardcore consumer hobbyist requirements. This is even further
You can run at ~20 tokens/second on a 512GB Mac Studio M3 Ultra: https://youtu.be/ufXZI6aqOU8?si=YGowQ3cSzHDpgv4z&t=197 IIRC the 512GB mac studio is about $10k