Live data from Hacker News

DSpark: Speculative decoding accelerates LLM inference [pdf]

github.com

221–230 of 393 posts

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#222
post #30

Earlier quoted context omitted.

Which is a good thing. Self-serving motives are more reliable than altruistic ones.

The world runs on incentives. Altruism/Self-serving are down stream of that. Wikipedia is altruistic, and serves humanity quite well.

Go read Max Stirner. True "Alturism" doesn't exist. It's all egoism, even if and especially if you think it's not.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#223
post #6

I’ve been using DeepSeek v4 pro for a month now in Kilo Code and its great. Fast, reliable, large context window and cheap as… Did 1,5B tokens this month and cost me 40usd (majority cached, but still).

Which provider? I went through 40 bucks on it on openrouter. It was not a lot of back and forth, context ended at around 300k, 15kloc output. I was using opencode, unsure if I can make the total token count visible.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#225

DeepSeek continues to not only push the boundaries but also publish these incredible papers explaining how they achieved their gains - something the American labs no longer do unfortunately. Chinese labs are doing the most interesting work in AI right now.

Probably because American AI companies are on the hook for quite a lot of investment money. I think they are trying to find the magical moat to justify their valuation. Revealing optimizations similar to these would pretty much reduce their competitive position.

> Probably because American AI companies are on the hook for quite a lot of investment money

That's a lot of words to say it's just capitalist greed.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#226

DeepSeek continues to not only push the boundaries but also publish these incredible papers explaining how they achieved their gains - something the American labs no longer do unfortunately. Chinese labs are doing the most interesting work in AI right now.

Google and Microsoft publish more than enough and American universities are publishing the science beyond DeepSeek's engineering. That fact that you don't know about them means you're not following the science only reading hacker news.

Google hasn’t published much in depth ML work since T5 (which was hugely influential at the time) - most Gemma releases are 1-3 page model card pdfs these days with no in depth analysis. Even TurboQuant is shaking out to have basically been a rehash of previous work without proper attribution. I do think Microsoft is doing some interesting things with smaller models but haven’t read much research, interested in any refs you might have to share!

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#227
post #6

I’ve been using DeepSeek v4 pro for a month now in Kilo Code and its great. Fast, reliable, large context window and cheap as… Did 1,5B tokens this month and cost me 40usd (majority cached, but still).

Have you compared Kilo to Pi or OpenCode? Those are the two I'm most familiar with but always looking for alternatives.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#228
post #30

Earlier quoted context omitted.

Which is a good thing. Self-serving motives are more reliable than altruistic ones.

The world runs on incentives. Altruism/Self-serving are down stream of that. Wikipedia is altruistic, and serves humanity quite well.

Is it though? A large number of people get to experience a lot of power over others because they moderate Wikipedia. That's certainly why some of them do it, just like on Reddit

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#229
post #112

Earlier quoted context omitted.

I don’t see how Anthropic is in a better position. They have a slight edge in model quality right at a time when we’re getting a taste of what cheap, “good enough” AI looks like. They don’t own their own compute. And their own arrogance and lies have alienated a huge chunk of their customer base and alerted everyone to the dangers of being dependent on them.

I personally think not owning their own compute is going to be an advantage. There is a meteor headed towards all this AI investment that I don't think has been properly accounted for and that is, what happens to all the existing hardware investments when NVidia's next architecture comes out. Blackwell (H100/H200) is the current generation. Rubin (R100, presumably R200) is the next and arrives soon. Now a lot of the…

Why do people who don't follow the prices of A100 talk like they know things about GPU pricing dynamics?

A100s are ~7 years old and going for more than 2 dollars an hour, significantly more expensive than even 2 years ago. This is because anything with 80gb of VRAM or more and made by Nvidia will have economically useful lifespans of like, 10 years.

I could see H100s getting 12 years.

Micheal Berry doesn't know shit about GPUs.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#230

I see a world soon where there’s an extremely wide variety of small models for speculative decoding, unique to use cases, companies, and even individuals.

You clearly didn't read the recent speculative decoding papers because it's been possible to use any model to speculate for any other model for awhile. They solved the tokenization problems that prevented this in the past.
Post reply on HN