Live data from Hacker News

DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

huggingface.co

341–350 of 485 posts

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#343
post #324

Earlier quoted context omitted.

I can’t think of a single major US company that is big internationally that is competing on price.

> I can’t think of a single major US company that is big internationally that is competing on price. All the clouds compete on price. Do you really think it is that differentiated? Google, Amazon and Microsoft all offer special deals to sign big companies up and globally too.

I worked inside AWS consulting department for 3 years (AWS ProServe) and now I work as a staff consultant for a 3rd AWS partner. I have been on enough sales calls, seen enough go to market training materials and flown out to customers sites to know how these things work. AWS has never tried to compete as the “low cost leader”. Marketing 101 says you never want to compete on price if you can avoid it.

Microsoft doesn’t compete on price. Their major competitive advantage is Big Enterprise is already big into Microsoft and it’s much easier to get them to come onto Azure. They compete on price only when it comes to making Windows workloads Bd SQL Server cheaper than running on other providers.

AWS is the default choice for legacy reasons and it definitely has services an offerings that Google doesn’t have. I have never once been on a sales call where the sales person emphasizes that AWS is cheaper.

As far as GCP, they are so bad at evterprise sales, we never really looked at them as serious competition.

Sure AWS will throw credits in for migrations and professional services both internally and for third party partners. But no CFO is going to look at just the short term credits.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#344
post #181

Earlier quoted context omitted.

If you're trying to build AI based applications you can and should compare the costs between vendor based solutions and hosting open models with your own hardware. On the hardware side you can run some benchmarks on the hardware (or use other people's benchmarks) and get an idea of the tokens/second you can get from the machine. Normalize this for your usage pattern (and do your best to implement batch processing whe…

Well the seemingly cheap comes with significantly degraded performance, particular for agentic use. Have you tried replacing Claude Code with some locally deployed model, say, on 4090 or 5090? I have. It is not usable.

Deepseek and Kimi both have great agentic performance

When used with crush/opencode they are close to Claude performance.

Nothing that runs on a 4090 would compete but Deepseek on openrouter is still 25x cheaper than claude

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#345
post #85

Earlier quoted context omitted.

I don't care if this kills Google and OpenAI. I hope it does, though I'm doubtful because distribution is important. You can't beat "ChatGPT" as a brand in laypeople's minds (unless perhaps you give them a massive "Temu: Shop Like A Billionaire" commercial campaign). Closed source AI is almost by design morphing into an industrial, infrastructure-heavy rocket science that commoners can't keep up with. The companies p…

I can’t think of a single company I’ve worked with as a consultant that I could convince to use DeepSeek because of its ties with China even if I explained that it was hosted on AWS and none of the information would go to China. Even when the technical people understood that, it would be too much of a political quagmire within their company when it became known to the higher ups. It just isn’t worth the political cap…

The average person has been programmed to be distrustful of open source in general, thinking it is inferior quality or in service of some ulterior motive

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#346
post #257

Earlier quoted context omitted.

For what it's worth, this is complete insanity when practically every mega enterprises' hardware is largely Made in China.

Enterprise hardware isn’t the issue. It’s the software. How much enterprise hardware is running with Chinese software? The US basically bans any hardware with Chinese software that can disrupt infrastructure.

Tons of routers, modems, embedded, are running Chinese software

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#347

Earlier quoted context omitted.

I can’t think of a single company I’ve worked with as a consultant that I could convince to use DeepSeek because of its ties with China even if I explained that it was hosted on AWS and none of the information would go to China. Even when the technical people understood that, it would be too much of a political quagmire within their company when it became known to the higher ups. It just isn’t worth the political cap…

[flagged]

Literally every time a Chinese model is discussed here we get this completely braindead take

There has never been a shred of evidence for security researchers, model analysis, benchmarks, etc that supports this.

It's a complete delusion in every sense.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#348

Earlier quoted context omitted.

>How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? According to Google (or someone at Google) no organization has moat on AI/LLM [1]. But that does not mean that it is not hugely profitable providing it as SaaS even you don't own the model or Model as a Service (MaaS). The extreme example is Amazon providing MongoDB API and services. Sure they have…

That quote from Google is 2.5 years old.

Have they said differently since?

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#349

Earlier quoted context omitted.

This is the rare earth minerals dumping all over again. Devalue to such a price as to make the market participants quit, so they can later have a strategic stranglehold on the supply. This is using open source in a bit of different spirit than the hacker ethos, and I am not sure how I feel about it. It is a kind of cheat on the fair market but at the same time it is also costly to China and its capital costs may beco…

The way we fund the AI bubble in the west could also be described as: "kind of cheat on the fair market". OpenAI has never made a single dime of profit.

Yeah and OpenAI's CPO was artificially commissioned as a Lt. Colonel in the US Army in conjunction with a $200M contract

Absurd to say Deepseek is CCP controlled while ignoring the govt connection here

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#350
post #308

Earlier quoted context omitted.

Yes. Though we don't know for sure whether that's because they actually have lower costs, or whether it's just the Chinese taxpayer being forced to serve us a treat.

Third party providers are still cheap though. The closed models are the ones where you can't see the real cost to running them.

Oh, I was mostly talking about the Chinese taxpayer footing the training bill.

You are right that we can directly observe the cost of inference for open models.

Post reply on HN