Live data from Hacker News

DSpark: Speculative decoding accelerates LLM inference [pdf]

github.com

351–360 of 393 posts

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#351

Earlier quoted context omitted.

Aren't they only open weights, not true open source?

The concept of open source doesn't really apply to AI models since their behavior is mostly controlled by the data they were trained on and the complex ways they are trained. Having the source code of the model by itself wouldn't help you. From a practical POV having all the training data, training infrastructure, and training know-how wouldn't help you either unless you could afford to spend the millions of dollars…

If (and that is a big if) the concept of open source doesn't apply, then the term shouldn't be coopted to mean something else though.

But even if I can't build it from source locally, being able to see what went into the model is an important part of what open source is about.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#352

Earlier quoted context omitted.

How very machiavellianist-libertarian of you. Don't even try to combine it with any notion of "leadership" then, however, since distillation is literally "copying the actual leader"

It really isn't, you can improve by distilling a weaker model

Self-distillation is also a technique.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#353

Earlier quoted context omitted.

If your moat is “please don’t copy my outputs”, you don’t have a moat. There is no such thing as a distillation “attack”.

How does it differ from pirating music or movies?

There is no intellectual property to “pirate”.

Model outputs don't qualify for copyright. They aren’t patented. They aren’t trade secrets - the companies sell them. They aren’t trademarks, obviously. They are nothing, actually.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#354

Earlier quoted context omitted.

You call yourself a philosopher in your profile.

I am also a fan of giving credit where it is due, and not giving it where it is not due. China is not an innovator. Perhaps they are only beginning to be, but historically, this is simply not the case, and yet distillation still falls squarely under "not innovating".

> historically, this is simply not the case

I wish people knew more history before using the term “historically”

The Chinese civilisation has been one of the most long living ones with plenty of “innovation” throughout history, both long ago and recently.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#355
post #314

Earlier quoted context omitted.

Given the MTP drafter is basically a separate model, keeping it separate makes more sense IMO. It's out of my wheelhouse but it seems like you could adjust the MTP drafter model separately from the main model, too. Ultimately though the real explanation, I think, is Google doesn't care since for their own purposes (in LiteRT-LM), they do bundle them. As far as I know, anyway.

Being grafted onto the main model reduces layer duplication that you’d otherwise have: at least for Step and Qwen 3.6

Step 2.7’s MTP seems broken (at least for ik_llama.cpp) where the draft model starts and ends in block 3 but ik_llama bails out looking for block 0 :(

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#356
post #341

Earlier quoted context omitted.

It's even worse than that. China publishes stacks upon stacks of policy documents in which they explain clearly what they will do and why. This includes why they do poverty alleviation and why they believe big monopolies that own everything are bad. But almost no western observers care to read those documents. Instead, western observers, including HN, speculate endlessly about China's intentions, and "it would be nai…

I had the most ironic rollercoaster ride thanks to your comment. I copied it into DeepSeek because I figured who's better to teach me about greatness of Chinese government policies if not the most popular Chinese LLM? Anyhow, it must have detected _something_ in your comment because Chinese censorship policy kicked in and DeepSeek refused to talk about it. Funny because I would wager the overall sentiment about China…

Chinese censorship rules are not about criticizing vs praising the state. They are about avoiding any kind of social controversy or collective action. They don't care when you talk about these topics in private or small circles, even criticism is fine, they just don't want them to spread far no matter whether it's positive or negative.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#357
post #341

Earlier quoted context omitted.

I had the most ironic rollercoaster ride thanks to your comment. I copied it into DeepSeek because I figured who's better to teach me about greatness of Chinese government policies if not the most popular Chinese LLM? Anyhow, it must have detected _something_ in your comment because Chinese censorship policy kicked in and DeepSeek refused to talk about it. Funny because I would wager the overall sentiment about China…

Chinese censorship rules are not about criticizing vs praising the state. They are about avoiding any kind of social controversy or collective action. They don't care when you talk about these topics in private or small circles, even criticism is fine, they just don't want them to spread far no matter whether it's positive or negative.

[dead]

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#358
post #314

Earlier quoted context omitted.

Being grafted onto the main model reduces layer duplication that you’d otherwise have: at least for Step and Qwen 3.6

Step 2.7’s MTP seems broken (at least for ik_llama.cpp) where the draft model starts and ends in block 3 but ik_llama bails out looking for block 0 :(

Aw that’s a shame; I’m running the official llama.cpp on my Spark-alike, and it works great now. Proper triple head too which is what it is trained on, gets me up to 35-40tk/s decode

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#359

Earlier quoted context omitted.

The concept of open source doesn't really apply to AI models since their behavior is mostly controlled by the data they were trained on and the complex ways they are trained. Having the source code of the model by itself wouldn't help you. From a practical POV having all the training data, training infrastructure, and training know-how wouldn't help you either unless you could afford to spend the millions of dollars…

If (and that is a big if) the concept of open source doesn't apply, then the term shouldn't be coopted to mean something else though. But even if I can't build it from source locally, being able to see what went into the model is an important part of what open source is about.

> If (and that is a big if) the concept of open source doesn't apply, then the term shouldn't be coopted to mean something else though.

Yes, but for whatever reason this usage seems to have stuck. Open weights is definitely a better name. I assume the reason "open source" has stuck is because you can download and use it for free, but "open source" was always intended to be about "free as in speech", not "free as in beer". That said, I remember when the term "open source" was invented, and it was always a bit different, more commercially aligned, than the goals of the FSF.

> But even if I can't build it from source locally, being able to see what went into the model is an important part of what open source is about.

True. Unfortunately LLMs have become such a big money and closed enterprise (the opposite of OpenAI and Anthropic's altruistic founding principles) that it's hard to see these commercial models releasing their training data, especially since this data is the closest thing they have to a moat other than the cost of training.

The most valuable training data right now seems to be "reasoning data", and the need for this at least may disappear as AI moves beyond pre-trained language models to smarter systems capable of learning for themselves, and that can actually reason, not need to parrot reasoning data.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#360
post #349

Earlier quoted context omitted.

Until recently, DeepSeek were self-financed (it was a spin-out from a hedge fund). They just raised ~50million RMB (US$7bn), and according to media [0] (which admittedly can be unreliable), the lead investors were: 1) The CEO himself 2) Tencent 3) CALT (the battery company) 4) NetEase (internet/media company) 5) JD.com (ecommerce) 6) Chinese investment firms What are they expecting in return? I'd say the same thing t…

If they’re expecting profits like the VC backed US companies how come they behave so differently?

The investment round only closed in the last few weeks, it would have had zero influence on anything up to & including DeepSeek V4.

Whether that now changes, who knows. It appears the CEO will remain the largest investor/shareholder, though, so I'm hoping not much changes.

Post reply on HN