Live data from Hacker News

DSpark: Speculative decoding accelerates LLM inference [pdf]

github.com

151–160 of 393 posts

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#151
post #85

Earlier quoted context omitted.

Probably because American AI companies are on the hook for quite a lot of investment money. I think they are trying to find the magical moat to justify their valuation. Revealing optimizations similar to these would pretty much reduce their competitive position.

Who is financing DeepSeek and what are they expecting in return?

I don't think this question would get to the reason. There could be one or two persons in charge who simply shape the culture of the company, including how much to publish.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#152

Earlier quoted context omitted.

If your moat is “please don’t copy my outputs”, you don’t have a moat. There is no such thing as a distillation “attack”.

How does it differ from pirating music or movies?

That when I pay for a model, the copyright of the output belongs to me. This is as work for hire as it gets.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#153
post #104
post #87

Earlier quoted context omitted.

They are self financed, the company that makes DeepSeek is a finance company that trades on the markets.

Even if they were fully self-financed, which isn’t the case, they would expect something in return.

Not everyone has the American “fuck you got mine” zero sum game attitude. Also they’re making some of the American and European AI companies look bad which they can leverage with their trades if they wanted to.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#154
post #24

Earlier quoted context omitted.

Chinese labs are also still behind, so they’re incentivized to collaborate and have no reason to do it in private. I suspect their tune will change if they ever take the lead..

> Chinese labs are also still behind, so they’re incentivized to collaborate and have no reason to do it in private. US labs in Google, Meta and SpaceX are not leading, none of them managed to build something on par with GLM 5.2. Care to explain to me why they still don't collaborate and still choose to do it in private?

Gemini 3.1 is still up there, though? If Google started to compete on price they could be very successful.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#155

Earlier quoted context omitted.

Publishing by necessity I wonder? American labs on the cutting edge pioneering the way forward, so Deepseek open sourcing what they’ve got is to help even the playing field. Hopefully the experts here can offer insight. The above is just my hunch and I’m not a specialist in this field.

Chinese papers and techniques have been very influential and copied by US labs. Multi-head Latent Attention (MLA), Multi-Token prediction, MoE architecture are some of the most famous examples.

MoE is from Google (Noam Shazeer)

MTP is from Meta

Another DeepSeek advance that the west are copying is DeepSeek Sparse Attention (DSA)

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#156
post #144

Earlier quoted context omitted.

Chinese labs are also still behind, so they’re incentivized to collaborate and have no reason to do it in private. I suspect their tune will change if they ever take the lead..

The question is also what game they're playing. Deepseek came out of a hedge fund. I think it's no coincidence that their publications tend to have a large impact on AI stock prices. Destroying the growth story of overvalued stocks is an interesting investment strategy. It's not even new. Shortsellers understandably get terrible rep from execs, but their actions are more often in the public interest than you'd think.…

> The question is also what game they're playing. Deepseek came out of a hedge fund. I think it's no coincidence that their publications tend to have a large impact on AI stock prices.

Its revealing that they always seem to publish after some big announcement by American AI companies. But regardless, this is one of the benefits of a duopoly.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#157

Earlier quoted context omitted.

If your moat is “please don’t copy my outputs”, you don’t have a moat. There is no such thing as a distillation “attack”.

How does it differ from pirating music or movies?

According to US AI labs, training on other people's output is fair use. So that's how.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#158
post #69

At this point why can't someone produce a fridge or container-sized AI appliance based on legacy chips (12nm)? I imagine this would cover 80% of corporate use cases where you need to "google-in-a-box" functionality. The state-of-the-art nanometer are impossible to achieve but if you have infinite solar energy during business hours does it really matter? Every company has a parking spot so this ASIC-like appliance cou…

The issue is that there are only so many fabs in the world that make memory. And if you want the good stuff, your easily going into 400 ~ 750b parameter models. That means at FP4 400 to 750GB memory. Did i mention there are only so many memory makers and they are all busy printing money with HBM memory? Intel is trying with Crescent Island, to make a 160GB GPU that uses LPDDR5X memory. HBM takes multiple times the re…

With MoE models like Deepseek’s and with multiple Crescent Island accelerators, the aggregate memory throughput actually doesn’t look that bad. Two Crescent Island gets roughly 1400GB/s and Deepseek-v4-flash with 13B parameters active nets roughly 100t/s which is decent for a small team or great for a single user.

More Crescent Island scale up, although not likely entirely linearly.

But all GPU inference work like this, it’s not specific to Intel. Just Intel promises more affordable cards with big memory so they’re attractive.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#159

DeepSeek continues to not only push the boundaries but also publish these incredible papers explaining how they achieved their gains - something the American labs no longer do unfortunately. Chinese labs are doing the most interesting work in AI right now.

Deepseek is commoditizing the performance gains US labs rely on to make their investors money.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#160

DeepSeek is, as I feel currently, the sole AI company which is actually trying to innovate rather than top mere benchmarks. Others like OpenAI, Anthropic and Google are mostly just competeing with each rather than keep innovating around the clock.

Qwen as well.
Post reply on HN