Live data from Hacker News

On DeepSeek and export controls

darioamodei.com

11–20 of 185 posts

Re: On DeepSeek and export controls

#11

It's shortsighted to believe that export controls are relevant. China will be able to manufacture as good or better chips than the US in the not too distant future anyway.

This is a choice between the lesser of two evils, or what could be called "drinking poison to quench thirst." Export controls might lead China to develop its own advanced chips, but without them, China will certainly gain powerful AI capabilities

Re: On DeepSeek and export controls

#12
> All of this is to say that DeepSeek-V3 is not a unique breakthrough or something that fundamentally changes the economics of LLM’s; it’s an expected point on an ongoing cost reduction curve. What’s different this time is that the company that was first to demonstrate the expected cost reductions was Chinese.

Says the CEO whose product [1] costs 15-50x times more. (This is not just the DeepSeek's API, but also 3p providers hosting the same model)

> DeepSeek does not "do for $6M5 what cost US AI companies billions". I can only speak for Anthropic, but Claude 3.5 Sonnet is a mid-sized model that cost a few $10M's to train

Ok, that's still at least 3-10x cost reduction (assuming "a few $10M" lowerbound of $20M). And for a model that he later implies is 2x larger than Sonnet. So that's 6-10x efficiency improvement. Nice!

> Since DeepSeek-V3 is worse than those US frontier models — let’s say by ~2x on the scaling curve.

What curve? Does he mean the simplistic performance / model params curve? That does not take into account that DeepSeek v3 is a MoE (can't compare MoE and dense param # in a naive way), nor the other architecture changes (KV compression, etc.).

Also, if Sonnet 3.5 is 2x smaller, then why is inference 15-50x more expensive than DeepSeek v3's? Does Anthropic not have good GPU engineers? Are they just running at insanely high margins? As a consumer I don't care how big your model is behind the scenes. I care about API costs or inference efficiency when hosting the model myself.

[1] Product that is mostly comparable and in some ways quite ahead.

Re: On DeepSeek and export controls

#13
post #4

Earlier quoted context omitted.

Anthropic is, according to themselves, using RLAIF... which is basically using LLM as a judge / reward model. So maybe he means that the models they use for RLAIF are not (much?) more expensive than Sonnet 3.5 (e.g. previous Sonnet or Haiku 3 :)).

Do you have a link to Anthropic saying they use RLAIF?

https://www.anthropic.com/research/constitutional-ai-harmles...

Re: On DeepSeek and export controls

#14
post #12

> All of this is to say that DeepSeek-V3 is not a unique breakthrough or something that fundamentally changes the economics of LLM’s; it’s an expected point on an ongoing cost reduction curve. What’s different this time is that the company that was first to demonstrate the expected cost reductions was Chinese. Says the CEO whose product [1] costs 15-50x times more. (This is not just the DeepSeek's API, but also 3p pr…

Where does he imply that it's 2x larger than Sonnet?

Re: On DeepSeek and export controls

#15
I just want to ask if when we talk about export controls on AI chips, is this going to create export controls on general consumer goods in the near future.

Given Moore’s law and the efficiency that will clearly come from optimizing chips for AI and competition increasing the amount of VRAM on these devices to run models locally.

Is creating an export regime today going to mean that in 3 or 4 years general smart phones and high end laptops are all going to be subject to export controls?

Keep in mind that at one time computers that consumed entire building are less powerful than my apple watch.

Re: On DeepSeek and export controls

#17
post #5

> Thus, in this world, the US and its allies might take a commanding and long-lasting lead on the global stage. This is just like Captain America giving a pre-battle speech to the Avengers—it's so inspiring! Hail Hydra!

Democratic USA must prevail against Communist China by having better models. edit: (Please read this in a sarcastic voice, I think it's a crazy idea!)

otherwise it will become to a multipolar world(or 'bipolar', said the article)

that would be horrible, a multipolar world? what a nightmare

Re: On DeepSeek and export controls

#18
post #5

> Thus, in this world, the US and its allies might take a commanding and long-lasting lead on the global stage. This is just like Captain America giving a pre-battle speech to the Avengers—it's so inspiring! Hail Hydra!

Democratic USA must prevail against Communist China by having better models. edit: (Please read this in a sarcastic voice, I think it's a crazy idea!)

OpenAI already announced they will be copying DeepSeek implementation, and that they will make sure not to share it with anyone else than themselves

Re: On DeepSeek and export controls

#19
post #17

Earlier quoted context omitted.

Democratic USA must prevail against Communist China by having better models. edit: (Please read this in a sarcastic voice, I think it's a crazy idea!)

otherwise it will become to a multipolar world(or 'bipolar', said the article) that would be horrible, a multipolar world? what a nightmare

As users, we would have more choices, more competition, and access to cheaper AI, more freedom (if DeepSeek stays open), higher quality models.

The same with chips, if no sanctions on exports to NVIDIA, we would have access to Huawei Ascend chips (used by DeepSeek to have lower runtime costs)

Re: On DeepSeek and export controls

#20
I'm kind of tired of the entire narrative that "US good, everyone else bad", as if only US deserves to hold powerful technology because it will be used for "the greater good", rather than employ it in military applications, as if Anthropic didn't partner with Palantir.

And it amazes me how many people can't seem to see past the "US is good, everyone else are bad" smoke mirror.

Post reply on HN