Live data from Hacker News

DSpark: Speculative decoding accelerates LLM inference [pdf]

github.com

131–140 of 393 posts

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#131
post #6

I’ve been using DeepSeek v4 pro for a month now in Kilo Code and its great. Fast, reliable, large context window and cheap as… Did 1,5B tokens this month and cost me 40usd (majority cached, but still).

I've been using omp with deepseek as my task and quicktask agents, and sonnet as everything else.

It's drastically reduced my AI spend. I went from spending $40/day to $10/day.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#132
post #126

Earlier quoted context omitted.

Chinese labs are also still behind, so they’re incentivized to collaborate and have no reason to do it in private. I suspect their tune will change if they ever take the lead..

Not everyone is motivated by greed

What do you think is the underlaying motivation?

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#133
post #85

Earlier quoted context omitted.

Probably because American AI companies are on the hook for quite a lot of investment money. I think they are trying to find the magical moat to justify their valuation. Revealing optimizations similar to these would pretty much reduce their competitive position.

Who is financing DeepSeek and what are they expecting in return?

IMHO to promote that China believes in free markets and making the technology available to all.

Which will likely help them bolster the sales of the MANY new AI chips in development/use in China to international markets. Dislodging Nvidia.

Kinda the opposite of what Jensen Huang (Nvidia) thinks US is doing: https://www.youtube.com/shorts/u3SY8nvjhQA

Edit: I'm a fan of deepseek and believe it's good to make the technology open/available. And do think that also help business - which I support as well.

Edit 2: No idea why I'm getting downvoted. That's also their official stance https://english.www.gov.cn/news/202601/08/content_WS695f1b55...

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#135

Earlier quoted context omitted.

If your moat is “please don’t copy my outputs”, you don’t have a moat. There is no such thing as a distillation “attack”.

How does it differ from pirating music or movies?

Machine-extruded text is not copyrightable, since there was no human creativity involved in producing it.

(and if you argue the US models do produce copyrighted works, then oooops - whose copyright is it huh?)

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#136

DeepSeek continues to not only push the boundaries but also publish these incredible papers explaining how they achieved their gains - something the American labs no longer do unfortunately. Chinese labs are doing the most interesting work in AI right now.

Sure, in part by "stealing" from American AI companies with Distillation attacks: https://yipzap.com/anthropic-accuses-alibaba-of-largest-ai-d...

US AI companies trained their own models on vast amounts of copyrighted and publicly available content without obtaining permission. There's no moral high ground here.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#137

DeepSeek continues to not only push the boundaries but also publish these incredible papers explaining how they achieved their gains - something the American labs no longer do unfortunately. Chinese labs are doing the most interesting work in AI right now.

Yep. It's about time western world realized Chinese are not the "very bad guys under dictatorship"

Honestly it's just a hierarchy difference between the two countries. In the US, tech/fin/military companies have the upper hand compared to the government (fragmented between 2 parties). Despite the sharades with Anthropic, Tech-fluencers are in control. Compared to china, the government (dictatorship) has more control over Tech companies (take any example from the past 10 years). For them, undermining the US AI supremacy is an objective, and releasing open weight models is the way, and I'm all for it.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#138

Earlier quoted context omitted.

If your moat is “please don’t copy my outputs”, you don’t have a moat. There is no such thing as a distillation “attack”.

How does it differ from pirating music or movies?

Ow my head.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#139
post #96

DeepSeek continues to not only push the boundaries but also publish these incredible papers explaining how they achieved their gains - something the American labs no longer do unfortunately. Chinese labs are doing the most interesting work in AI right now.

Its because our culture worships pieces of paper the government tells us is worth something.

Nope, people seek it out because government tells them to pay taxes _or else_.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#140
post #9

This is just one of many papers DeepSeek have released to be able to serve models at extremely cheap prices, unlike the others taking on >$100B+ of debt in building data centers for the same thing. > As with V4-Flash, we treat this point as an indication that DSpark sustains useful throughput under an interactivity target that the baseline cannot efficiently support. At matched system capacities, DSpark delivers 57%…

For so long American companies have operated under the assumption that servers are cheaper than developers, and that was used to justify all sorts of inefficient practices.

The last year has shown that’s not true anymore (even for web servers).

Post reply on HN