Earlier quoted context omitted.
Probably because American AI companies are on the hook for quite a lot of investment money. I think they are trying to find the magical moat to justify their valuation. Revealing optimizations similar to these would pretty much reduce their competitive position.
Who is financing DeepSeek and what are they expecting in return?
DSpark: Speculative decoding accelerates LLM inference [pdf]
151–160 of 393 posts
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#152Earlier quoted context omitted.
If your moat is “please don’t copy my outputs”, you don’t have a moat. There is no such thing as a distillation “attack”.
How does it differ from pirating music or movies?
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#153Earlier quoted context omitted.
They are self financed, the company that makes DeepSeek is a finance company that trades on the markets.
Even if they were fully self-financed, which isn’t the case, they would expect something in return.
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#154Earlier quoted context omitted.
Chinese labs are also still behind, so they’re incentivized to collaborate and have no reason to do it in private. I suspect their tune will change if they ever take the lead..
> Chinese labs are also still behind, so they’re incentivized to collaborate and have no reason to do it in private. US labs in Google, Meta and SpaceX are not leading, none of them managed to build something on par with GLM 5.2. Care to explain to me why they still don't collaborate and still choose to do it in private?
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#155Earlier quoted context omitted.
Publishing by necessity I wonder? American labs on the cutting edge pioneering the way forward, so Deepseek open sourcing what they’ve got is to help even the playing field. Hopefully the experts here can offer insight. The above is just my hunch and I’m not a specialist in this field.
Chinese papers and techniques have been very influential and copied by US labs. Multi-head Latent Attention (MLA), Multi-Token prediction, MoE architecture are some of the most famous examples.
MTP is from Meta
Another DeepSeek advance that the west are copying is DeepSeek Sparse Attention (DSA)
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#156Earlier quoted context omitted.
Chinese labs are also still behind, so they’re incentivized to collaborate and have no reason to do it in private. I suspect their tune will change if they ever take the lead..
The question is also what game they're playing. Deepseek came out of a hedge fund. I think it's no coincidence that their publications tend to have a large impact on AI stock prices. Destroying the growth story of overvalued stocks is an interesting investment strategy. It's not even new. Shortsellers understandably get terrible rep from execs, but their actions are more often in the public interest than you'd think.…
Its revealing that they always seem to publish after some big announcement by American AI companies. But regardless, this is one of the benefits of a duopoly.
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#157Earlier quoted context omitted.
If your moat is “please don’t copy my outputs”, you don’t have a moat. There is no such thing as a distillation “attack”.
How does it differ from pirating music or movies?
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#158At this point why can't someone produce a fridge or container-sized AI appliance based on legacy chips (12nm)? I imagine this would cover 80% of corporate use cases where you need to "google-in-a-box" functionality. The state-of-the-art nanometer are impossible to achieve but if you have infinite solar energy during business hours does it really matter? Every company has a parking spot so this ASIC-like appliance cou…
The issue is that there are only so many fabs in the world that make memory. And if you want the good stuff, your easily going into 400 ~ 750b parameter models. That means at FP4 400 to 750GB memory. Did i mention there are only so many memory makers and they are all busy printing money with HBM memory? Intel is trying with Crescent Island, to make a 160GB GPU that uses LPDDR5X memory. HBM takes multiple times the re…
More Crescent Island scale up, although not likely entirely linearly.
But all GPU inference work like this, it’s not specific to Intel. Just Intel promises more affordable cards with big memory so they’re attractive.
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#159DeepSeek continues to not only push the boundaries but also publish these incredible papers explaining how they achieved their gains - something the American labs no longer do unfortunately. Chinese labs are doing the most interesting work in AI right now.
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#160DeepSeek is, as I feel currently, the sole AI company which is actually trying to innovate rather than top mere benchmarks. Others like OpenAI, Anthropic and Google are mostly just competeing with each rather than keep innovating around the clock.