Live data from Hacker News

DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

github.com

71–80 of 218 posts

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#71

Earlier quoted context omitted.

amusingly ive been working on ultra sparse llm inference/ training/ model design because nature loaths a dense graph/matrix and cause i think it shoukd be possible. i actually stood up a 20-25 percent faster than sota causal fast attention kernel yesterday, will be standing up cuda/metal/armv8 kernels too and thats gonna be fun. i genuinely think these models should be like 0.1 percent sparse for same capabilities we…

Look forwarding your future releases

i definitely will be doing some drop of some faster attention kernels in the next few weeks.

like i can do all sorts of memory layout of tensors/matrices etc tricks that if you dont have the abstractions for it would just never happen. so i can optimize the kernel flops

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#72
post #39

Curious what the fundamental limit on Huawei's capacity is. China has shown if nothing else they know how to scale when they want to. If it came down to just building more of what they know how to do, it would be happening. Is there more to it?

Huawei chips need advanced 3d packaging in order to keep up. Since the process is too complex the yield is still bad.

Not to mention there is a lot of demand from various factors, not deepseek only. Huawei itself is a major consumer.

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#73

Here's something I really don't understand: If as alleged Chinese open weight models are catching up with US anyway, and the performance is near US frontier model level but Chinese can do it with a fraction of cost, and eventually AI model will be commodified, wouldn't that means that the billion or even trillion dollars that US labs spend have only diminishing returns and the lead is only temporary? So why Deepseek…

there's a lot of propganda from these state backed enterprises. I think the fraction of the cost label is debatable given the evidence of mass gpu smuggling through third parties like Singapore which China can't exactly openly admit to. Unless of course we're talking about distilling, which is probably a lot cheaper than training a model from scratch (there's also the fact that labour is still relatively cheap in Chi…

All that and so what? Fact is the Chinese have several near peer models, they've released the weights and they are widely available.

You want to sue them or something?

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#74

I think the way to parse the current title "DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]" is that there was a leak that DeepSeek will pause fundraising because they perceive there is a compute gap with the US. I am also guessing that the majority of the people who read this title will think that DeepSeek is pausing this fundraising because some comments they made about the co…

All the Chinese reporting I see point to the second (majority) interpretation. Liang being furious about his private investor talk leaked online is the news here. e.g. https://x.com/_FORAB/status/2081034500101017616?s=20

Those could be subsequent developments, but that's not what the linked transcript was about. The transcript was a discussion of the DeepSeek founder (Liang Wenfeng) with investors, and he does not mention any leaks, or any frustration. He simply says that he is constrained by the supply of cards, and he has no problem of getting funding, but has no reason to raise further funding because he can't transform the cash into cards.

  > There is certainly no shortage of funds or resources --- in fact, all these are readily available [...]

  > Within our financial capacity, it's undoubtedly true that the more cards are always better. Our current strategy is to purchase as many cards as possible at a reasonable price --- exactly how many we can afford after using this funding round. The spending pace isn't predetermined; we'll buy whatever is available as long as prices remain competitive. In fact, I'd consider that a positive outcome if we spend the entire amount within six months. [...]

  > In reality, spending such a large sum is no easy task: you can't obtain enough cards, they're hard to come by [...]

  > Therefore, our only concern is whether we can obtain enough cards. If converting all funds into cards were feasible, we would undoubtedly do so without hesitation and are even willing to pay a premium for this benefit --- it's simply to cost effective. Even after paying the premium, however, achieving this goal remains challenging.

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#75
post #47

Here's something I really don't understand: If as alleged Chinese open weight models are catching up with US anyway, and the performance is near US frontier model level but Chinese can do it with a fraction of cost, and eventually AI model will be commodified, wouldn't that means that the billion or even trillion dollars that US labs spend have only diminishing returns and the lead is only temporary? So why Deepseek…

U.S. policymakers believe that even if the gap is small—like six months to a year—whoever reaches AGI first (whatever that means) could gain such an overwhelming advantage over their perceived adversary that it would effectively kneecap them. (You can look at the kinds of things they mention—cyber, WMDs—to get a sense of what they mean.) Jensen Huang disagrees and has said AI is a marathon.

It’s kind of true but also kind of silly.

True in that frontier models do have the capability to outperform all other models, but silly because AGI self improvement is itself an iterative process that takes a lot of compute.

So you can imagine a world where all the frontier labs achieve AGI but in order to keep their AGI ahead of other AGIs they have to use more and more compute until all the compute is going to self improvement and there is nothing left for other tasks.

That is just a silly scenario so I think when AGI is around we will still have bottlenecks that force it to grow at a moderate rate instead of asymptomatically.

AGI first mover advantage implies that there is no such bottlenecks.

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#76

Here's something I really don't understand: If as alleged Chinese open weight models are catching up with US anyway, and the performance is near US frontier model level but Chinese can do it with a fraction of cost, and eventually AI model will be commodified, wouldn't that means that the billion or even trillion dollars that US labs spend have only diminishing returns and the lead is only temporary? So why Deepseek…

Deepseek is funded by their hedge fund, high flyer. They intentionally cap their token prices to basically recoup server costs. The meeting transcript describes it as a moral commitment, that they don’t care about trends like image and video generation, and world model “hype”. They only care about reasoning, chain of thought and continuous learning.

One word: Focus.

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#77

Here's something I really don't understand: If as alleged Chinese open weight models are catching up with US anyway, and the performance is near US frontier model level but Chinese can do it with a fraction of cost, and eventually AI model will be commodified, wouldn't that means that the billion or even trillion dollars that US labs spend have only diminishing returns and the lead is only temporary? So why Deepseek…

there's a lot of propganda from these state backed enterprises. I think the fraction of the cost label is debatable given the evidence of mass gpu smuggling through third parties like Singapore which China can't exactly openly admit to. Unless of course we're talking about distilling, which is probably a lot cheaper than training a model from scratch (there's also the fact that labour is still relatively cheap in Chi…

Is there really a valid basis for claims like "state backed enterprises"? My understanding is that's no different than claiming that datacentres in America are "state backed" since they get things like big breaks on property taxes.

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#78
post #59
post #57

Earlier quoted context omitted.

If that is the case, it means one thing only - US labs don't have moat whatsoever and their expectation to have trillion dollar valuation is just laughable.

The moat is the compute.

The compute for training, or for inference?

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#80

Here's something I really don't understand: If as alleged Chinese open weight models are catching up with US anyway, and the performance is near US frontier model level but Chinese can do it with a fraction of cost, and eventually AI model will be commodified, wouldn't that means that the billion or even trillion dollars that US labs spend have only diminishing returns and the lead is only temporary? So why Deepseek…

My understanding after reading Liang’s comments during the investment meeting is that Liang firmly believes in AGI and he bets everything to reach goal. Once it reaches AGI, the game would flip totally. How he didn’t paint it out, and with the potential severe impact on the labor and consumer market, the true economic impact is difficult to predict. Liang is more like religious about this goal.

He also admits that it’s still a long way to it and along the way you have to recoup some money, too. But that is not their main motive, because focus too much on this short term goal will lower their probability of AGI success and it’s trivial to what AGI can bring. Liang stressed on restraining and emphasized that it’s part of their culture.

Thus, they continue invest in AI because they believe in breakthrough and not just being better.

Post reply on HN