Live data from Hacker News

On DeepSeek and export controls

darioamodei.com

1–10 of 185 posts

Re: On DeepSeek and export controls

#2
> DeepSeek does not "do for $6M what cost US AI companies billions". I can only speak for Anthropic, but Claude 3.5 Sonnet is a mid-sized model that cost a few $10M's to train (I won't give an exact number). Also, 3.5 Sonnet was not trained in any way that involved a larger or more expensive model (contrary to some rumors).

Wow!

Re: On DeepSeek and export controls

#3

> DeepSeek does not "do for $6M what cost US AI companies billions". I can only speak for Anthropic, but Claude 3.5 Sonnet is a mid-sized model that cost a few $10M's to train (I won't give an exact number). Also, 3.5 Sonnet was not trained in any way that involved a larger or more expensive model (contrary to some rumors). Wow!

The narrative is running away in every direction.

There are popular infographics floating around today which explain to the layman that deepseek invented MoE.

Re: On DeepSeek and export controls

#4

> DeepSeek does not "do for $6M what cost US AI companies billions". I can only speak for Anthropic, but Claude 3.5 Sonnet is a mid-sized model that cost a few $10M's to train (I won't give an exact number). Also, 3.5 Sonnet was not trained in any way that involved a larger or more expensive model (contrary to some rumors). Wow!

Anthropic is, according to themselves, using RLAIF... which is basically using LLM as a judge / reward model. So maybe he means that the models they use for RLAIF are not (much?) more expensive than Sonnet 3.5 (e.g. previous Sonnet or Haiku 3 :)).

Re: On DeepSeek and export controls

#5
> Thus, in this world, the US and its allies might take a commanding and long-lasting lead on the global stage.

This is just like Captain America giving a pre-battle speech to the Avengers—it's so inspiring! Hail Hydra!

Re: On DeepSeek and export controls

#7
At least they're all coming out of the woodwork and start telling more details about their own training runs, costs, efficiency rates and so on. Interesting to see how an open-weights model could force their hand like that.

Re: On DeepSeek and export controls

#8
post #4

> DeepSeek does not "do for $6M what cost US AI companies billions". I can only speak for Anthropic, but Claude 3.5 Sonnet is a mid-sized model that cost a few $10M's to train (I won't give an exact number). Also, 3.5 Sonnet was not trained in any way that involved a larger or more expensive model (contrary to some rumors). Wow!

Anthropic is, according to themselves, using RLAIF... which is basically using LLM as a judge / reward model. So maybe he means that the models they use for RLAIF are not (much?) more expensive than Sonnet 3.5 (e.g. previous Sonnet or Haiku 3 :)).

Do you have a link to Anthropic saying they use RLAIF?

Re: On DeepSeek and export controls

#9
post #5

> Thus, in this world, the US and its allies might take a commanding and long-lasting lead on the global stage. This is just like Captain America giving a pre-battle speech to the Avengers—it's so inspiring! Hail Hydra!

Democratic USA must prevail against Communist China by having better models.

edit: (Please read this in a sarcastic voice, I think it's a crazy idea!)

Re: On DeepSeek and export controls

#10
Unless the author makes a compelling case about why AI breaks the MAD status-quo between nuclear powers, I will assume that their appeals to NATSEC are an attempt to artificially create moat for their company.
Post reply on HN