Live data from Hacker News

DSpark: Speculative decoding accelerates LLM inference [pdf]

github.com

71–80 of 393 posts

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#72
post #46
post #19

Earlier quoted context omitted.

Exactly. They did not have to open up their research up and this is what happens when smart researchers are forced to squeeze performance gains out of existing hardware. They don't have TPUs or access to the latest Vera Rubin GPUs either to get performance gains for free. All of the optimizations Deepseek have done are in software and it goes down to the PTX assembly level. Compared to Anthropic who are celebrating i…

Anthropic almost certainly also has optimized software down to the assembly level, considering this take-home interview challenge they published: https://github.com/anthropics/original_performance_takehome/... which is all about instruction-level performance optimizations. That they don't prioritize UI fixes just means they consider other things more important.

that's pretty silly to use as a measure of what they do internally

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#73

Presumably this has been in production for a while, and is one of the reasons they were able to dramatically lower prices a month ago?

Yes. Section 5 talks about real-world deployment: 5.1: "The DSpark draft models are co-deployed with the preview versions of DeepSeek-V4-Flash and DeepSeek-V4-Pro"; 5.4: "MTP-1 represents the former production setup, having been superseded by DSpark two weeks following the DeepSeek-V4-preview release."

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#74
post #46
post #19

Earlier quoted context omitted.

Exactly. They did not have to open up their research up and this is what happens when smart researchers are forced to squeeze performance gains out of existing hardware. They don't have TPUs or access to the latest Vera Rubin GPUs either to get performance gains for free. All of the optimizations Deepseek have done are in software and it goes down to the PTX assembly level. Compared to Anthropic who are celebrating i…

Anthropic almost certainly also has optimized software down to the assembly level, considering this take-home interview challenge they published: https://github.com/anthropics/original_performance_takehome/... which is all about instruction-level performance optimizations. That they don't prioritize UI fixes just means they consider other things more important.

Unlikely: that product is written completely by AI, of which they are not lacking.

More likely is that an AI generated codename is impossible to fix by humans, and SOTA was not able to figure it out until now.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#75
post #66
post #50

Earlier quoted context omitted.

Chinese companies (and labs) operate in conjunction with the CCP so whatever they're doing, it's because it's Chinese state policy. What became clear when DeepSeek came onto the scene was that China was seeking to commoditize LLMs. They consider it an issue of national security not to be beholden to US tech companies when it comes to AI. And I, for one, fully endorse this policy. Another data point on this is the bla…

> Another data point on this is the black market for Claude tokens in China [1]. The chat logs themselves are a commodity to train models. anyone with IQ higher than 130 (thus qualified for actual AI R&D) would be questioning something obvious here - if they are already doing such dodgy stuff with the aim to maximize profits, why would those resellers have large amount of logs with actual American model responses to…

[deleted]

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#76
post #53
post #30

Earlier quoted context omitted.

Which is a good thing. Self-serving motives are more reliable than altruistic ones.

You mean more predictable, not more reliable.

I don't think so. I can confidently predict that altruism will give you a very unreliable income stream in the vast majority of cases.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#77
post #69

At this point why can't someone produce a fridge or container-sized AI appliance based on legacy chips (12nm)? I imagine this would cover 80% of corporate use cases where you need to "google-in-a-box" functionality. The state-of-the-art nanometer are impossible to achieve but if you have infinite solar energy during business hours does it really matter? Every company has a parking spot so this ASIC-like appliance cou…

Nvidia is already selling exactly this I think, not sure when it's expected to ship

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#79
post #29

I am wondering if this is why they can offer their pro model at ~1/4th of the price compared to the other providers offering the same model, and if other providers will be able to do the same in a short timeframe.

I have been heavily using DeepSeek V4 Pro at Max for a month now and I would say it is 100x cheaper. If I pay for Claude I will hit that limit so fast I am always waiting 5 hours. Using the frontier models at Kilo I go through dollars while doing the same thing via DeepSeek it is pennies.

I believe the comment you replied to was talking about the cost on providers like OpenCode vs Deepseek API. Deepseek API is even cheaper than the other providers for the same deepseek models.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#80
post #30

Earlier quoted context omitted.

Chinese labs are also still behind, so they’re incentivized to collaborate and have no reason to do it in private. I suspect their tune will change if they ever take the lead..

Which is a good thing. Self-serving motives are more reliable than altruistic ones.

The world runs on incentives. Altruism/Self-serving are down stream of that.

Wikipedia is altruistic, and serves humanity quite well.

Post reply on HN