Live data from Hacker News

DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

github.com

201–210 of 218 posts

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#201
post #190

Earlier quoted context omitted.

The real question is whether it's easier to improve the software side instead. There are likely a lot more optimizations possible in terms of model architecture, and if there is a compute bottleneck, then it's going to put a lot of pressure on Chinese labs to address the problem using more efficient designs.

And one cool trick is that if one develops step function more efficient training, one can still “open source” the model without revealing the training techniques.

exactly

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#202
post #16

Earlier quoted context omitted.

Wow, that is fascinating, I didn't realize China was now blocking foreign chips, lol. It's not a definitive indicator, but I feel that doesn't bode well for US dominance in this area -- when your competitor thinks they'd be helping _you_ by using your resources, that's not great.

They've achieved self-sufficiency in >14nm chips in remarkable timing. Unfortunately for DeepSeek, it's the And even if Huawei's Ascend 910C can compete with NVIDIA's H200, CUDA is still a large moat

China is already making it's own chips. The one most commonly discussed is the Huawei Ascend 950 (which exists and is in production). NVIDIA say the 950 is about equivalent to an NVIDIA H100, and DeepSeek in this leaked meeting are saying that about 4 950 cards are equivalent to 1 NVIDIA "B" one (Blackwell B200 I think).

Apparently production is constrained though, with DeepSeek saying they were only able to get an allocation of 16,000 950 cards. They have plenty of money, but are limited by the number of cards available to buy, and therefore the size of models they can train. They mentioned having a 20,000 GPU cluster. Other companies like Kimi, with a ~3T param model, clearly have a lot more compute, probably all NVIDIA.

DeepSeek have their own "TileLang" software, which sounds a bit like the Triton kernel compiler, and isolates them from the diffences between NVIDIA and Huawei chips - they are deliberately avoiding any CUDA dependency.

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#203
post #39

Curious what the fundamental limit on Huawei's capacity is. China has shown if nothing else they know how to scale when they want to. If it came down to just building more of what they know how to do, it would be happening. Is there more to it?

The production capacity constraint seems to come from SMIC who make the Ascend processors for Huawei. Huawei's memory comes from CXMT who seem to have plenty of capacity, with Apple looking to buy memory from them. Huawei then combines processors and memory into chiplets similar to what NVIDIA does with their GPUs.

The reason SMIC are capacity constrained is at least in part because they've been blocked from buying ASML's EUV machines, and are therefore having to make do with previous generation lower resolution DUV machines. These DUV machines can be coaxed into making surprisingly competitive 5-7nm chips, but at the expense of using many more production steps ("multi patterning") which limits productivity.

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#204
post #199

Earlier quoted context omitted.

>Unless I’ve missed some advancement? nah they're still just statistical token predictors based on their training data, solving hundred year old math conjectures one day, only just given the formulation; strictly benchmarkmaxxing with all guardrails turned off by deciding to look up the answers to their benchmark questions by zero daying their airgap, hopping over to the third party that hosts the answers, zero dayin…

I get that you're being cheeky here. Yes, LLM capabilities have expanded. We might be working with different definitions of "Artificial General Intelligence" here, for which there is no agreed-upon formal definition[1]. I was thinking of the "thinking, reasoning, maybe feeling" kind when I wrote my comment. But if you're thinking along the "really good at technical tasks" definition, sure, maybe. [1]: https://en.wiki…

cheeky? I was super serious! There's a list of criteria right in the section you hot-linked and these things have never come close to doing any of the things on that list.

>for which there is no agreed-upon formal definition

we all agree that the definition is not whatever this is.

Also, even if these things ever did seem to think, reason, or feel, we all agree that they still don't really though.

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#206
post #32
post #10

Earlier quoted context omitted.

Trump reversed course on the NVIDIA ban. It's now China that is blocking their companies from buying NVIDIA chips. So the shell entities would be to get around Chinese, not USian restrictions

That's not true. First there is still a licensing and quota scheme on the US side for the H200s. Secondly China blocked them for use in inferencing. Thirdly Chinese companies don't want them for training because newer chips are more cost effective.

He says he needs 200,000 GB300s or Huawei Ascend 950s. Now, the GB300 is much faster. So, probably he's just paying a little lip service to the 950 because of political pressure and would prefer the former. It's a moot point though because Hauwei can't manufacture with a low enough defect rate currently.

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#207

Everything in this transcript reads so very different from what megalomaniacs in charge of Anthropic/OAI have to say

Did you also read the transcripts of meetings of Anthropic/OpenAI's investors? Maybe you should read its IPO Filing. Since the doc isn't available at the moment, may be try SpaceX's one to see how an official doc of a company of another "megalomaniac" looks like, especially the section "CAUTIONARY STATEMENT REGARDING FORWARD-LOOKING STATEMENTS" https://www.sec.gov/Archives/edgar/data/1181412/000162828026...

[deleted]

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#208

Earlier quoted context omitted.

is this similar to Meta starting to rent out own compute as they cant seem to do much with it and monetising it is much better ROI?...

Sorry to barge in here. I couldn't find a good place to place https://github.com/demo-zexuan/liang-wenfeng-investor-meetin... The current link is 404, can mods update to above, detach, make sticky? (No response from mods, understandable)

I'm not sure I understand the question but I've replaced the top link (https://github.com/demo-zexuan/liang-wenfeng-investor-meetin...), which was 404ing, with the link in your comment here. Does that help?

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#209

Earlier quoted context omitted.

It's much easier to distill a model than create one from scratch. Part of the reason the open source model factories have been able to keep par with the frontier model factories is that they distill the frontier models, not recreate something as good from scratch.

> It's much easier to distill a model than create one from scratch. Part of the reason the open source model factories have been able to keep par with the frontier model factories is that they distill the frontier models, not recreate something as good from scratch. I think the "why" was "why would the US companies have models that can't be distilled?", not "why does distillation work"?

I suppose it would be due to intense research and development. Obviously anyone could come upon these models if they exist, so it isn't necessarily a given that frontier labs would do it first, but probably the odds should be given to the groups with the biggest pull and the largest budgets

Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]

#210
post #208

Earlier quoted context omitted.

Sorry to barge in here. I couldn't find a good place to place https://github.com/demo-zexuan/liang-wenfeng-investor-meetin... The current link is 404, can mods update to above, detach, make sticky? (No response from mods, understandable)

I'm not sure I understand the question but I've replaced the top link ( https://github.com/demo-zexuan/liang-wenfeng-investor-meetin... ), which was 404ing, with the link in your comment here. Does that help?

Yes, thank you!

It's now dropped off the FrontPage but I suppose I will just repost at some point with a less controversial title

Post reply on HN