DeepSeek continues to not only push the boundaries but also publish these incredible papers explaining how they achieved their gains - something the American labs no longer do unfortunately. Chinese labs are doing the most interesting work in AI right now.
Probably because American AI companies are on the hook for quite a lot of investment money. I think they are trying to find the magical moat to justify their valuation. Revealing optimizations similar to these would pretty much reduce their competitive position.
DSpark: Speculative decoding accelerates LLM inference [pdf]
41–50 of 393 posts
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#42DeepSeek continues to not only push the boundaries but also publish these incredible papers explaining how they achieved their gains - something the American labs no longer do unfortunately. Chinese labs are doing the most interesting work in AI right now.
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#43Earlier quoted context omitted.
Probably because American AI companies are on the hook for quite a lot of investment money. I think they are trying to find the magical moat to justify their valuation. Revealing optimizations similar to these would pretty much reduce their competitive position.
I seriously am far from fear mongering and doomsday mentality, but I just can't see how OpenAI and Anthropic can have a successful IPO if the quality gap between the free and paid continues to narrow like that...
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#44DeepSeek continues to not only push the boundaries but also publish these incredible papers explaining how they achieved their gains - something the American labs no longer do unfortunately. Chinese labs are doing the most interesting work in AI right now.
Publishing by necessity I wonder? American labs on the cutting edge pioneering the way forward, so Deepseek open sourcing what they’ve got is to help even the playing field. Hopefully the experts here can offer insight. The above is just my hunch and I’m not a specialist in this field.
Multi-head Latent Attention (MLA), Multi-Token prediction, MoE architecture are some of the most famous examples.
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#45Must be wonderful to be on the board of OpenAi et al & their PE investors whilst China keeps blowing up these mines under their feet lmao. Luckily Korean pension funds will buy all the trash as usual but goddamn you gotta start moving quick or you are gonna need some serious AGI to show you how to offload those bonds
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#46DeepSeek continues to not only push the boundaries but also publish these incredible papers explaining how they achieved their gains - something the American labs no longer do unfortunately. Chinese labs are doing the most interesting work in AI right now.
Exactly. They did not have to open up their research up and this is what happens when smart researchers are forced to squeeze performance gains out of existing hardware. They don't have TPUs or access to the latest Vera Rubin GPUs either to get performance gains for free. All of the optimizations Deepseek have done are in software and it goes down to the PTX assembly level. Compared to Anthropic who are celebrating i…
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#47Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#48Earlier quoted context omitted.
Which is a good thing. Self-serving motives are more reliable than altruistic ones.
Very interesting take
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#49Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#50DeepSeek continues to not only push the boundaries but also publish these incredible papers explaining how they achieved their gains - something the American labs no longer do unfortunately. Chinese labs are doing the most interesting work in AI right now.
What became clear when DeepSeek came onto the scene was that China was seeking to commoditize LLMs. They consider it an issue of national security not to be beholden to US tech companies when it comes to AI. And I, for one, fully endorse this policy.
Another data point on this is the black market for Claude tokens in China [1]. The chat logs themselves are a commodity to train models.
I believe that OpenAI in particular is a bet on a trillion dollar pot of gold that doesn't exist. Google, Microsoft, Amazon and Meta will all be fine. Anthropic is in a far better position than OpenAI (IMHO) but if DeepSeek or some other Chinese open weight model gets as good at coding, they're in real trouble too.