Live data from Hacker News

DSpark: Speculative decoding accelerates LLM inference [pdf]

github.com

51–60 of 393 posts

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#51
post #37
post #29

I am wondering if this is why they can offer their pro model at ~1/4th of the price compared to the other providers offering the same model, and if other providers will be able to do the same in a short timeframe.

It'd presumably help a lot, but also when you use their endpoint they get more training data.

US labs do it too.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#53
post #30

Earlier quoted context omitted.

Chinese labs are also still behind, so they’re incentivized to collaborate and have no reason to do it in private. I suspect their tune will change if they ever take the lead..

Which is a good thing. Self-serving motives are more reliable than altruistic ones.

You mean more predictable, not more reliable.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#54
post #39
post #30

Earlier quoted context omitted.

Which is a good thing. Self-serving motives are more reliable than altruistic ones.

Very interesting take

It's a standard take since it is how markets tend to work. They aren't powered by altruism, it is a big system for turning greed into good results. We don't have all this stuff because people suddenly woke up one morning and decided to be nice.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#55
post #41

Earlier quoted context omitted.

I seriously am far from fear mongering and doomsday mentality, but I just can't see how OpenAI and Anthropic can have a successful IPO if the quality gap between the free and paid continues to narrow like that...

[flagged]

[flagged]

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#57
post #24

Earlier quoted context omitted.

Chinese labs are also still behind, so they’re incentivized to collaborate and have no reason to do it in private. I suspect their tune will change if they ever take the lead..

> Chinese labs are also still behind, so they’re incentivized to collaborate and have no reason to do it in private. US labs in Google, Meta and SpaceX are not leading, none of them managed to build something on par with GLM 5.2. Care to explain to me why they still don't collaborate and still choose to do it in private?

Google at least still releases open source models to the public.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#58
post #6

I’ve been using DeepSeek v4 pro for a month now in Kilo Code and its great. Fast, reliable, large context window and cheap as… Did 1,5B tokens this month and cost me 40usd (majority cached, but still).

Is there a way to see how many tokes one does with claude code (pro)?

It's in the JSONs in ~/.claude, but last 30 days only I think. You can have the model analyze history. So for correct history you'd need to run history analysis on a cron job or something. Kinda hacky.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#59

Earlier quoted context omitted.

Probably because American AI companies are on the hook for quite a lot of investment money. I think they are trying to find the magical moat to justify their valuation. Revealing optimizations similar to these would pretty much reduce their competitive position.

[flagged]

I’m think its in our best interests to lever these american ai companies to exhibit at least some degree of freedom and transparency anyway we can…

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#60
post #40
post #37

Earlier quoted context omitted.

It'd presumably help a lot, but also when you use their endpoint they get more training data.

This applies to every provider. OpenAI seems to be the worst hoarder.

actually you can buy inference on third party providers that serve deepseek v4 pro with zero data retention (ZDR).
Post reply on HN