Earlier quoted context omitted.
The real question is whether it's easier to improve the software side instead. There are likely a lot more optimizations possible in terms of model architecture, and if there is a compute bottleneck, then it's going to put a lot of pressure on Chinese labs to address the problem using more efficient designs.
And one cool trick is that if one develops step function more efficient training, one can still “open source” the model without revealing the training techniques.
DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]
201–210 of 218 posts
Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]
#202Earlier quoted context omitted.
Wow, that is fascinating, I didn't realize China was now blocking foreign chips, lol. It's not a definitive indicator, but I feel that doesn't bode well for US dominance in this area -- when your competitor thinks they'd be helping _you_ by using your resources, that's not great.
They've achieved self-sufficiency in >14nm chips in remarkable timing. Unfortunately for DeepSeek, it's the And even if Huawei's Ascend 910C can compete with NVIDIA's H200, CUDA is still a large moat
Apparently production is constrained though, with DeepSeek saying they were only able to get an allocation of 16,000 950 cards. They have plenty of money, but are limited by the number of cards available to buy, and therefore the size of models they can train. They mentioned having a 20,000 GPU cluster. Other companies like Kimi, with a ~3T param model, clearly have a lot more compute, probably all NVIDIA.
DeepSeek have their own "TileLang" software, which sounds a bit like the Triton kernel compiler, and isolates them from the diffences between NVIDIA and Huawei chips - they are deliberately avoiding any CUDA dependency.
Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]
#203Curious what the fundamental limit on Huawei's capacity is. China has shown if nothing else they know how to scale when they want to. If it came down to just building more of what they know how to do, it would be happening. Is there more to it?
The reason SMIC are capacity constrained is at least in part because they've been blocked from buying ASML's EUV machines, and are therefore having to make do with previous generation lower resolution DUV machines. These DUV machines can be coaxed into making surprisingly competitive 5-7nm chips, but at the expense of using many more production steps ("multi patterning") which limits productivity.
Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]
#204Earlier quoted context omitted.
>Unless I’ve missed some advancement? nah they're still just statistical token predictors based on their training data, solving hundred year old math conjectures one day, only just given the formulation; strictly benchmarkmaxxing with all guardrails turned off by deciding to look up the answers to their benchmark questions by zero daying their airgap, hopping over to the third party that hosts the answers, zero dayin…
I get that you're being cheeky here. Yes, LLM capabilities have expanded. We might be working with different definitions of "Artificial General Intelligence" here, for which there is no agreed-upon formal definition[1]. I was thinking of the "thinking, reasoning, maybe feeling" kind when I wrote my comment. But if you're thinking along the "really good at technical tasks" definition, sure, maybe. [1]: https://en.wiki…
>for which there is no agreed-upon formal definition
we all agree that the definition is not whatever this is.
Also, even if these things ever did seem to think, reason, or feel, we all agree that they still don't really though.
Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]
#205Is there a mirror? Getting 404...
Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]
#206Earlier quoted context omitted.
Trump reversed course on the NVIDIA ban. It's now China that is blocking their companies from buying NVIDIA chips. So the shell entities would be to get around Chinese, not USian restrictions
That's not true. First there is still a licensing and quota scheme on the US side for the H200s. Secondly China blocked them for use in inferencing. Thirdly Chinese companies don't want them for training because newer chips are more cost effective.
Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]
#207Everything in this transcript reads so very different from what megalomaniacs in charge of Anthropic/OAI have to say
Did you also read the transcripts of meetings of Anthropic/OpenAI's investors? Maybe you should read its IPO Filing. Since the doc isn't available at the moment, may be try SpaceX's one to see how an official doc of a company of another "megalomaniac" looks like, especially the section "CAUTIONARY STATEMENT REGARDING FORWARD-LOOKING STATEMENTS" https://www.sec.gov/Archives/edgar/data/1181412/000162828026...
Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]
#208Earlier quoted context omitted.
is this similar to Meta starting to rent out own compute as they cant seem to do much with it and monetising it is much better ROI?...
Sorry to barge in here. I couldn't find a good place to place https://github.com/demo-zexuan/liang-wenfeng-investor-meetin... The current link is 404, can mods update to above, detach, make sticky? (No response from mods, understandable)
Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]
#209Earlier quoted context omitted.
It's much easier to distill a model than create one from scratch. Part of the reason the open source model factories have been able to keep par with the frontier model factories is that they distill the frontier models, not recreate something as good from scratch.
> It's much easier to distill a model than create one from scratch. Part of the reason the open source model factories have been able to keep par with the frontier model factories is that they distill the frontier models, not recreate something as good from scratch. I think the "why" was "why would the US companies have models that can't be distilled?", not "why does distillation work"?
Re: DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]
#210Earlier quoted context omitted.
Sorry to barge in here. I couldn't find a good place to place https://github.com/demo-zexuan/liang-wenfeng-investor-meetin... The current link is 404, can mods update to above, detach, make sticky? (No response from mods, understandable)
I'm not sure I understand the question but I've replaced the top link ( https://github.com/demo-zexuan/liang-wenfeng-investor-meetin... ), which was 404ing, with the link in your comment here. Does that help?
It's now dropped off the FrontPage but I suppose I will just repost at some point with a less controversial title