Earlier quoted context omitted.
Which is a good thing. Self-serving motives are more reliable than altruistic ones.
Very interesting take
DSpark: Speculative decoding accelerates LLM inference [pdf]
71–80 of 393 posts
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#72Earlier quoted context omitted.
Exactly. They did not have to open up their research up and this is what happens when smart researchers are forced to squeeze performance gains out of existing hardware. They don't have TPUs or access to the latest Vera Rubin GPUs either to get performance gains for free. All of the optimizations Deepseek have done are in software and it goes down to the PTX assembly level. Compared to Anthropic who are celebrating i…
Anthropic almost certainly also has optimized software down to the assembly level, considering this take-home interview challenge they published: https://github.com/anthropics/original_performance_takehome/... which is all about instruction-level performance optimizations. That they don't prioritize UI fixes just means they consider other things more important.
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#73Presumably this has been in production for a while, and is one of the reasons they were able to dramatically lower prices a month ago?
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#74Earlier quoted context omitted.
Exactly. They did not have to open up their research up and this is what happens when smart researchers are forced to squeeze performance gains out of existing hardware. They don't have TPUs or access to the latest Vera Rubin GPUs either to get performance gains for free. All of the optimizations Deepseek have done are in software and it goes down to the PTX assembly level. Compared to Anthropic who are celebrating i…
Anthropic almost certainly also has optimized software down to the assembly level, considering this take-home interview challenge they published: https://github.com/anthropics/original_performance_takehome/... which is all about instruction-level performance optimizations. That they don't prioritize UI fixes just means they consider other things more important.
More likely is that an AI generated codename is impossible to fix by humans, and SOTA was not able to figure it out until now.
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#75Earlier quoted context omitted.
Chinese companies (and labs) operate in conjunction with the CCP so whatever they're doing, it's because it's Chinese state policy. What became clear when DeepSeek came onto the scene was that China was seeking to commoditize LLMs. They consider it an issue of national security not to be beholden to US tech companies when it comes to AI. And I, for one, fully endorse this policy. Another data point on this is the bla…
> Another data point on this is the black market for Claude tokens in China [1]. The chat logs themselves are a commodity to train models. anyone with IQ higher than 130 (thus qualified for actual AI R&D) would be questioning something obvious here - if they are already doing such dodgy stuff with the aim to maximize profits, why would those resellers have large amount of logs with actual American model responses to…
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#76Earlier quoted context omitted.
Which is a good thing. Self-serving motives are more reliable than altruistic ones.
You mean more predictable, not more reliable.
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#77At this point why can't someone produce a fridge or container-sized AI appliance based on legacy chips (12nm)? I imagine this would cover 80% of corporate use cases where you need to "google-in-a-box" functionality. The state-of-the-art nanometer are impossible to achieve but if you have infinite solar energy during business hours does it really matter? Every company has a parking spot so this ASIC-like appliance cou…
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#78Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#79I am wondering if this is why they can offer their pro model at ~1/4th of the price compared to the other providers offering the same model, and if other providers will be able to do the same in a short timeframe.
I have been heavily using DeepSeek V4 Pro at Max for a month now and I would say it is 100x cheaper. If I pay for Claude I will hit that limit so fast I am always waiting 5 hours. Using the frontier models at Kilo I go through dollars while doing the same thing via DeepSeek it is pennies.
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#80Earlier quoted context omitted.
Chinese labs are also still behind, so they’re incentivized to collaborate and have no reason to do it in private. I suspect their tune will change if they ever take the lead..
Which is a good thing. Self-serving motives are more reliable than altruistic ones.
Wikipedia is altruistic, and serves humanity quite well.