I am wondering if this is why they can offer their pro model at ~1/4th of the price compared to the other providers offering the same model, and if other providers will be able to do the same in a short timeframe.
It'd presumably help a lot, but also when you use their endpoint they get more training data.
DSpark: Speculative decoding accelerates LLM inference [pdf]
51–60 of 393 posts
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#52Would love to see these numbers reproduced on consumer GPUs, not just A100s.
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#53Earlier quoted context omitted.
Chinese labs are also still behind, so they’re incentivized to collaborate and have no reason to do it in private. I suspect their tune will change if they ever take the lead..
Which is a good thing. Self-serving motives are more reliable than altruistic ones.
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#54Earlier quoted context omitted.
Which is a good thing. Self-serving motives are more reliable than altruistic ones.
Very interesting take
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#55Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#56Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#57Earlier quoted context omitted.
Chinese labs are also still behind, so they’re incentivized to collaborate and have no reason to do it in private. I suspect their tune will change if they ever take the lead..
> Chinese labs are also still behind, so they’re incentivized to collaborate and have no reason to do it in private. US labs in Google, Meta and SpaceX are not leading, none of them managed to build something on par with GLM 5.2. Care to explain to me why they still don't collaborate and still choose to do it in private?
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#58I’ve been using DeepSeek v4 pro for a month now in Kilo Code and its great. Fast, reliable, large context window and cheap as… Did 1,5B tokens this month and cost me 40usd (majority cached, but still).
Is there a way to see how many tokes one does with claude code (pro)?
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#59Earlier quoted context omitted.
Probably because American AI companies are on the hook for quite a lot of investment money. I think they are trying to find the magical moat to justify their valuation. Revealing optimizations similar to these would pretty much reduce their competitive position.
[flagged]
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#60Earlier quoted context omitted.
It'd presumably help a lot, but also when you use their endpoint they get more training data.
This applies to every provider. OpenAI seems to be the worst hoarder.