Live data from Hacker News

DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

github.com

81–90 of 218 posts

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#81
post #70

Earlier quoted context omitted.

Their yields on high performance chips that could do training is really bad, and they aren’t getting more of the outdated ASML machines that they could use to scale up even with bad yields. It will still take China a few years or a decade to build out the tech needed to fab high performance chips economically on their own.

Can't they use GlobalFoundries, Intel or TSMC in addition to their own fabs?

Chinese companies are effectively embargoed from using them for <14nm.

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#82
post #79

The repository was force-pushed so the link doesn't work anymore, but the file is still available at: https://github.com/demo-zexuan/liang-wenfeng-investor-meetin...

Thank you. These documents are a particularly valuable insight into the kind of thinking going on at DeepSeek. The part about how inference should be priced at a level that's enough to return capex in 10 months, but no higher, is really interesting. Liang Wenfeng simply has different motivations than we're used to over here in the West.

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#84

Earlier quoted context omitted.

there's a lot of propganda from these state backed enterprises. I think the fraction of the cost label is debatable given the evidence of mass gpu smuggling through third parties like Singapore which China can't exactly openly admit to. Unless of course we're talking about distilling, which is probably a lot cheaper than training a model from scratch (there's also the fact that labour is still relatively cheap in Chi…

All that and so what? Fact is the Chinese have several near peer models, they've released the weights and they are widely available. You want to sue them or something?

> You want to sue them or something?

I mean, if this were an american company vs an american company, i think it would be a long drawn out civil case and brought before the Supreme Court (I still this is ultimately will be brought before the supreme court). It could also be argued frontier models are far more important to national security than most military programs, even versus next gen fighter jets.

The fact that Alibaba stock, which is also listed on the NYSE, barely budged after Anthropic made these claims imo tells me that the market doesn't think that a lone american company could go after these companies by themselves. Alibaba denied and there's not much they can do alone, I mean would the CCP allow Alibaba go through a discovery process of a normal civil trial? It might have to be the US feds that bring up a case.

I think it could be argued that if Alibaba and other China companies want access to US capital markets for something so vital for national security, there should be some ground rules, but we will eventually need the Supreme court to settle whether or not this state enterprise distilling constitutes IP theft (at the very least it is a breach of contract). The fact that they are widely available doesn't really matter (i mean pirated content is widely available, it's ultimately about how the court rules on distilling).

based on this HN comment and associated article https://news.ycombinator.com/item?id=48977128#48985989 I still have yet to see a China open weight model beat any of the frontier models, they always almost there yet never quite there, which seems to be evidence of distilling (although I'm open to be proven wrong).

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#85
post #47

Here's something I really don't understand: If as alleged Chinese open weight models are catching up with US anyway, and the performance is near US frontier model level but Chinese can do it with a fraction of cost, and eventually AI model will be commodified, wouldn't that means that the billion or even trillion dollars that US labs spend have only diminishing returns and the lead is only temporary? So why Deepseek…

U.S. policymakers believe that even if the gap is small—like six months to a year—whoever reaches AGI first (whatever that means) could gain such an overwhelming advantage over their perceived adversary that it would effectively kneecap them. (You can look at the kinds of things they mention—cyber, WMDs—to get a sense of what they mean.) Jensen Huang disagrees and has said AI is a marathon.

> U.S. policymakers believe that even if the gap is small—like six months to a year—whoever reaches AGI first...

From past experience, AGI was never seriously discussed in these kinds of conversations beyond thought experiments, and was basically humoring SBF, Daniela Amodei, and the other EA types (some deep believers, but some who I felt were cynically using it as a way to preempt competition back when OpenAI and Google were the behemoths).

The big worry is applications of AI in C4ISR, OffSec, loitering munitions, Disinfo/social media botting (notice the recent shift towards identification on social media ;)), and other sorts of DefenseTech adjacent usecases.

The second worry is that an AI race turns into an infra buildout race, and HPC is extremely dual use, especially in the simulations space because of the NPT, the CTBT, and the PTBT.

The AGI-pilled people aren't the ones to worry about - it's the people who understand the limits of models and how to integrate with cyberphysical applications.

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#86
post #25

Earlier quoted context omitted.

Wow, that is fascinating, I didn't realize China was now blocking foreign chips, lol. It's not a definitive indicator, but I feel that doesn't bode well for US dominance in this area -- when your competitor thinks they'd be helping _you_ by using your resources, that's not great.

I don't know that it's clear that the motivation is that it's "helping" their competitors directly. Maybe the motivation is "if we rely on these, then the next time a US president arbitrarily decides to block us from buying them, we won't have the infrastructure already in place to be able to work around it". It seems more betting on a shorter-term cost with less uncertainty in the long term rather than a shorter-ter…

This was the Chinese government burning the boats[0]. Beyond the symbolism and ensuring everyone's commitment, they want to direct the firehose of AI money towards a local champion (Huawei). Money, and experience bourne of being in the trenches developing, manufacturing, deploying and debugging on actual workloads will supercharge how quickly Huawei get good, compared to when they had to fairly compete against a well-resourced Nvidia

0. Though it's not absolute - there's still the Singapore-based clusters loop-hole, that the Chinese government may choose the degree to which it turns a blind eye to, if progress is slow.

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#87
post #39

Curious what the fundamental limit on Huawei's capacity is. China has shown if nothing else they know how to scale when they want to. If it came down to just building more of what they know how to do, it would be happening. Is there more to it?

Their yields on high performance chips that could do training is really bad, and they aren’t getting more of the outdated ASML machines that they could use to scale up even with bad yields. It will still take China a few years or a decade to build out the tech needed to fab high performance chips economically on their own.

They can already make chips economically. Yields are worse than TSCM but good enough since they no longer need to pay the "Qualcomm tax". Huawei's phones are profitable. The issue is rather that Chinese capacity comes from a low quantity, and scaling capacity while simultaneously indigineousizing parts takes time — years. Foreign capacity was too good and too cheap so they never succeeded in scaling capacity, because the demand for Chinese fabs wasn't there. Now the demand for domestic capacity is there and they're scaling like crazy, like 100-200% growth per year. But demand still far outpaces capacity. Fabs are hard to build. You need many more years of 200% growth to even approach the demand.

And yes, I made use of emdash. It's a legit grammar tool. Sue me.

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#88
>> As you can understand, during V3 training, NVIDIA GPUs were still used, but the NVIDIA ecosystem was no longer employed.

Ironic that these large LLMs are eroding Nividia's moat. In the next paragraph he talks about Nvidia digging its own grave. I wonder if Nividia is aware of this and the frequent release cycle is a response to this development ?

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#89

Earlier quoted context omitted.

there's a lot of propganda from these state backed enterprises. I think the fraction of the cost label is debatable given the evidence of mass gpu smuggling through third parties like Singapore which China can't exactly openly admit to. Unless of course we're talking about distilling, which is probably a lot cheaper than training a model from scratch (there's also the fact that labour is still relatively cheap in Chi…

Is there really a valid basis for claims like "state backed enterprises"? My understanding is that's no different than claiming that datacentres in America are "state backed" since they get things like big breaks on property taxes.

Ownership and corporate governance is much more state-led via GGIFs as well as mandated party oversight depending on the size of company.

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#90
post #25

Earlier quoted context omitted.

I don't know that it's clear that the motivation is that it's "helping" their competitors directly. Maybe the motivation is "if we rely on these, then the next time a US president arbitrarily decides to block us from buying them, we won't have the infrastructure already in place to be able to work around it". It seems more betting on a shorter-term cost with less uncertainty in the long term rather than a shorter-ter…

This was the Chinese government burning the boats[0]. Beyond the symbolism and ensuring everyone's commitment, they want to direct the firehose of AI money towards a local champion (Huawei). Money, and experience bourne of being in the trenches developing, manufacturing, deploying and debugging on actual workloads will supercharge how quickly Huawei get good, compared to when they had to fairly compete against a well…

The money and resources aren't going to Huawei exclusively, but to a whole range of companies. This selection changes over time as new companies show promise, or previously selected companies prove themselves to be incompetent.

For example there are like 6 different technical tracks of EUV development. They're not committing to one technical direction, it's exploring all of them.

Post reply on HN