DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
341–350 of 485 posts
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#342Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#343Earlier quoted context omitted.
I can’t think of a single major US company that is big internationally that is competing on price.
> I can’t think of a single major US company that is big internationally that is competing on price. All the clouds compete on price. Do you really think it is that differentiated? Google, Amazon and Microsoft all offer special deals to sign big companies up and globally too.
Microsoft doesn’t compete on price. Their major competitive advantage is Big Enterprise is already big into Microsoft and it’s much easier to get them to come onto Azure. They compete on price only when it comes to making Windows workloads Bd SQL Server cheaper than running on other providers.
AWS is the default choice for legacy reasons and it definitely has services an offerings that Google doesn’t have. I have never once been on a sales call where the sales person emphasizes that AWS is cheaper.
As far as GCP, they are so bad at evterprise sales, we never really looked at them as serious competition.
Sure AWS will throw credits in for migrations and professional services both internally and for third party partners. But no CFO is going to look at just the short term credits.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#344Earlier quoted context omitted.
If you're trying to build AI based applications you can and should compare the costs between vendor based solutions and hosting open models with your own hardware. On the hardware side you can run some benchmarks on the hardware (or use other people's benchmarks) and get an idea of the tokens/second you can get from the machine. Normalize this for your usage pattern (and do your best to implement batch processing whe…
Well the seemingly cheap comes with significantly degraded performance, particular for agentic use. Have you tried replacing Claude Code with some locally deployed model, say, on 4090 or 5090? I have. It is not usable.
When used with crush/opencode they are close to Claude performance.
Nothing that runs on a 4090 would compete but Deepseek on openrouter is still 25x cheaper than claude
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#345Earlier quoted context omitted.
I don't care if this kills Google and OpenAI. I hope it does, though I'm doubtful because distribution is important. You can't beat "ChatGPT" as a brand in laypeople's minds (unless perhaps you give them a massive "Temu: Shop Like A Billionaire" commercial campaign). Closed source AI is almost by design morphing into an industrial, infrastructure-heavy rocket science that commoners can't keep up with. The companies p…
I can’t think of a single company I’ve worked with as a consultant that I could convince to use DeepSeek because of its ties with China even if I explained that it was hosted on AWS and none of the information would go to China. Even when the technical people understood that, it would be too much of a political quagmire within their company when it became known to the higher ups. It just isn’t worth the political cap…
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#346Earlier quoted context omitted.
For what it's worth, this is complete insanity when practically every mega enterprises' hardware is largely Made in China.
Enterprise hardware isn’t the issue. It’s the software. How much enterprise hardware is running with Chinese software? The US basically bans any hardware with Chinese software that can disrupt infrastructure.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#347Earlier quoted context omitted.
I can’t think of a single company I’ve worked with as a consultant that I could convince to use DeepSeek because of its ties with China even if I explained that it was hosted on AWS and none of the information would go to China. Even when the technical people understood that, it would be too much of a political quagmire within their company when it became known to the higher ups. It just isn’t worth the political cap…
[flagged]
There has never been a shred of evidence for security researchers, model analysis, benchmarks, etc that supports this.
It's a complete delusion in every sense.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#348Earlier quoted context omitted.
>How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? According to Google (or someone at Google) no organization has moat on AI/LLM [1]. But that does not mean that it is not hugely profitable providing it as SaaS even you don't own the model or Model as a Service (MaaS). The extreme example is Amazon providing MongoDB API and services. Sure they have…
That quote from Google is 2.5 years old.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#349Earlier quoted context omitted.
This is the rare earth minerals dumping all over again. Devalue to such a price as to make the market participants quit, so they can later have a strategic stranglehold on the supply. This is using open source in a bit of different spirit than the hacker ethos, and I am not sure how I feel about it. It is a kind of cheat on the fair market but at the same time it is also costly to China and its capital costs may beco…
The way we fund the AI bubble in the west could also be described as: "kind of cheat on the fair market". OpenAI has never made a single dime of profit.
Absurd to say Deepseek is CCP controlled while ignoring the govt connection here
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#350Earlier quoted context omitted.
Yes. Though we don't know for sure whether that's because they actually have lower costs, or whether it's just the Chinese taxpayer being forced to serve us a treat.
Third party providers are still cheap though. The closed models are the ones where you can't see the real cost to running them.
You are right that we can directly observe the cost of inference for open models.