Live data from Hacker News

DSpark: Speculative decoding accelerates LLM inference [pdf]

github.com

61–70 of 393 posts

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#61
post #29

I am wondering if this is why they can offer their pro model at ~1/4th of the price compared to the other providers offering the same model, and if other providers will be able to do the same in a short timeframe.

I have been heavily using DeepSeek V4 Pro at Max for a month now and I would say it is 100x cheaper. If I pay for Claude I will hit that limit so fast I am always waiting 5 hours. Using the frontier models at Kilo I go through dollars while doing the same thing via DeepSeek it is pennies.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#62
The hugging face models are already up and seem to be the original models with the speculative decoding module built in which is very cool:

Flash: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-DSpark

Pro: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-DSpark

Excited to see if this makes it into DwarfStar for local inference, have been using the flash model extensively since the 2-bit quants were made available by antirez.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#63

Earlier quoted context omitted.

Probably because American AI companies are on the hook for quite a lot of investment money. I think they are trying to find the magical moat to justify their valuation. Revealing optimizations similar to these would pretty much reduce their competitive position.

Chinese labs are also still behind, so they’re incentivized to collaborate and have no reason to do it in private. I suspect their tune will change if they ever take the lead..

Projection is a funny thing. It causes people to misread situations all the time. Southern slaveowners feared violent retribution from freed slaves, for example [1]. It was pure projection and said more about the South than it did the slaves. The reality was there was no violent retribution. It was the opposite where the former slaveowners continued to inflict violence on the formerly enslaved.

I say this because we see the same thing used as an argument against China. "If they overtake us, they'll do imperialism (like us)." Again, it says more about us than them.

A better reading (IMHO) Of the situation is that China believes that AI shouldn't be used simply to mint a few more trillionaires but the benefits should be shared with society. Why do I say this? Because we now have 70+ years of China doing exactly that. The transformation in China all the way from rural villages to Tier 1 cities has been utterly astounding. China has lifted ~800M people out of extreme poverty.

In some ways we're at a similar point to the late 1990s and 2000s when Microsoft execs complained that Linux, being free, destroyed intellectual property value. Linux should be a perfect example of how people can and do act altruistically, or at least not in a way to bait-and-switch to enrich themselves.

[1]: https://www.reddit.com/r/AskHistory/comments/1d26grm/in_the_...

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#64

Earlier quoted context omitted.

[flagged]

This is incorrect binary thinking. Them releasing open source can be good, but that does not commit you to think that china or chinese companies are saints. There are many shades of grey here and one does not exclude the other (nor include it).

Are you reading the comments?

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#65

Earlier quoted context omitted.

Publishing by necessity I wonder? American labs on the cutting edge pioneering the way forward, so Deepseek open sourcing what they’ve got is to help even the playing field. Hopefully the experts here can offer insight. The above is just my hunch and I’m not a specialist in this field.

Wouldn’t that just help the American labs anyway though? Or do they assume they’ve actually already figured this stuff out and kept it secret?

From what I gather, the Chinese are behind, but a lot of their research amounts to scrappy, clever discoveries in how to use more novel technologies (for Qwen and Deepseek, its mixture of expert models, that can do inference using a portion of the model at a time). The chinese also distill information from American models, so there’s that.

The American companies, from my impression don’t involve themselves with such lowly “hacks” because they have so much money to just push forward with doing everything on big heavy models that run on the most cutting edge nvidia chips that they can, the moment, kinda sorta get on demand (I say that in some degree of jest).

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#66
post #50

DeepSeek continues to not only push the boundaries but also publish these incredible papers explaining how they achieved their gains - something the American labs no longer do unfortunately. Chinese labs are doing the most interesting work in AI right now.

Chinese companies (and labs) operate in conjunction with the CCP so whatever they're doing, it's because it's Chinese state policy. What became clear when DeepSeek came onto the scene was that China was seeking to commoditize LLMs. They consider it an issue of national security not to be beholden to US tech companies when it comes to AI. And I, for one, fully endorse this policy. Another data point on this is the bla…

> Another data point on this is the black market for Claude tokens in China [1]. The chat logs themselves are a commodity to train models.

anyone with IQ higher than 130 (thus qualified for actual AI R&D) would be questioning something obvious here -

if they are already doing such dodgy stuff with the aim to maximize profits, why would those resellers have large amount of logs with actual American model responses to sell to those AI labs in the first place. shouldn't they just post train & customize some leading Chinese open source models to pretend to be Opus or GPT for the vast majority of their users (as classified by some models) who don't know much about expected Opus behaviours & not skilled enough to tell the differences?

that is actually the interesting bit not covered in your censored version of the story line, it is also what happens on the ground. your censored version of the story implies that those dodgy resellers using stolen credit cards, pooling accounts with stolen IDs and illegally selling very personal logs would somehow be honest enough to spend extra $ to ensure their victims (aka paying users) can actually use real Opus and GPT. LOL

dude, you failed this IQ test miserably.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#67
post #60
post #40

Earlier quoted context omitted.

This applies to every provider. OpenAI seems to be the worst hoarder.

actually you can buy inference on third party providers that serve deepseek v4 pro with zero data retention (ZDR).

Only reliable way to have zero data retention is to self-host.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#68
These companies providing tokens, whether SOTA or not, that want to IPO are so fucked as time goes on.

Can't sell their SOTA models, only slightly better than the open source models for the models they can sell, cost 20x to 50x for good models, a TAM that consists almost solely of developers, with no customer of theirs actually boasting increased profits as a result of AI...

I fear their time to IPO may have passed.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#69
At this point why can't someone produce a fridge or container-sized AI appliance based on legacy chips (12nm)? I imagine this would cover 80% of corporate use cases where you need to "google-in-a-box" functionality.

The state-of-the-art nanometer are impossible to achieve but if you have infinite solar energy during business hours does it really matter? Every company has a parking spot so this ASIC-like appliance could be as big as a shipping container.

If it could just run recent open models for a handful of users it would be such a nobrainer to buy.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#70
post #39
post #30

Earlier quoted context omitted.

Which is a good thing. Self-serving motives are more reliable than altruistic ones.

Very interesting take

I don't understand what is interesting about it: it's the default.

Markets don't run on altruism.

Post reply on HN