Earlier quoted context omitted.
Extremely interesting comment, thank you. Got some links where I can download this source material? I don't read or speak the language, but will try interrogating it with an LLM
The fifth book is on Amazon. https://www.amazon.com/XI-JINPING-GOVERNANCE-CHINA-V/dp/7119... It's already an English translation. For something shorter, you can see Arnaud Bertrand's recent review. https://arnaudbertrand.substack.com/p/the-book-the-west-refu... The review is behind a paywall, but not expensive. If you want to read policy documents directly (primary source), try the State Council / Chinese government…
DSpark: Speculative decoding accelerates LLM inference [pdf]
231–240 of 393 posts
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#232Earlier quoted context omitted.
The question is also what game they're playing. Deepseek came out of a hedge fund. I think it's no coincidence that their publications tend to have a large impact on AI stock prices. Destroying the growth story of overvalued stocks is an interesting investment strategy. It's not even new. Shortsellers understandably get terrible rep from execs, but their actions are more often in the public interest than you'd think.…
[flagged]
Chinese companies understand this and they're treating models as shared infrastructure akin to Linux. The money is going to be in customization niches. Companies will charge to tune models for specific use cases and charge support for that. There's also going to be money at the bottom for hardware vendors making chips and memory. But the middle tier of generic LLMs is seeing involution where there's relentless competition driving profits towards the bottom.
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#233Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#234Earlier quoted context omitted.
The question is also what game they're playing. Deepseek came out of a hedge fund. I think it's no coincidence that their publications tend to have a large impact on AI stock prices. Destroying the growth story of overvalued stocks is an interesting investment strategy. It's not even new. Shortsellers understandably get terrible rep from execs, but their actions are more often in the public interest than you'd think.…
> The question is also what game they're playing. Deepseek came out of a hedge fund. I think it's no coincidence that their publications tend to have a large impact on AI stock prices. Its revealing that they always seem to publish after some big announcement by American AI companies. But regardless, this is one of the benefits of a duopoly.
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#235Earlier quoted context omitted.
It used to be the case that NSA hired the majority of all math graduates in the US, and were assumed to be years ahead in cryptography. Yet in the 90s, it became clear that they no longer were that - among other things, the cipher of the notorious Clipper chip was broken, and we can rule out that it was made weak on purpose because the whole point of Clipper was that they had a backdoor. So, despite hiring the cream…
Everyone in this thread is getting distracted by nationalism, but you hit the nail on the head. In this case for whatever reason the Chinese AI industry is collaborative and the American AI industry is not. This will result in the Chinese companies making progress faster. Full stop. This isn't a judgement on the merits of either system, only an observation of likely results.
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#236Earlier quoted context omitted.
From what I gather, the Chinese are behind, but a lot of their research amounts to scrappy, clever discoveries in how to use more novel technologies (for Qwen and Deepseek, its mixture of expert models, that can do inference using a portion of the model at a time). The chinese also distill information from American models, so there’s that. The American companies, from my impression don’t involve themselves with such…
The American companies would love to develop these 'hacks' because it would make them more money, something they are in existential need of right now. They don't develop them because they don't collaborate publicly anymore. Where would the whole industry be if Google never allowed publishing the transformers paper? It's not a coincidence that the American AI industry grew fastest in capability when it was the most op…
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#237Earlier quoted context omitted.
The question is also what game they're playing. Deepseek came out of a hedge fund. I think it's no coincidence that their publications tend to have a large impact on AI stock prices. Destroying the growth story of overvalued stocks is an interesting investment strategy. It's not even new. Shortsellers understandably get terrible rep from execs, but their actions are more often in the public interest than you'd think.…
[flagged]
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#238Earlier quoted context omitted.
Which is a good thing. Self-serving motives are more reliable than altruistic ones.
The world runs on incentives. Altruism/Self-serving are down stream of that. Wikipedia is altruistic, and serves humanity quite well.
[†] Another problem with altruism: we don't all agree on whether a goal is altruistic, and what's altruistic in the enactor's eyes might not be in yours. Curating a fountain of human knowledge like Wikipedia? Probably altruistic. Protecting humanity from itself by installing your company as the stewards of frontier LLMs? Not so altruistic in my view.
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#239I’ve been using DeepSeek v4 pro for a month now in Kilo Code and its great. Fast, reliable, large context window and cheap as… Did 1,5B tokens this month and cost me 40usd (majority cached, but still).
Which provider? I went through 40 bucks on it on openrouter. It was not a lot of back and forth, context ended at around 300k, 15kloc output. I was using opencode, unsure if I can make the total token count visible.
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#240DeepSeek is, as I feel currently, the sole AI company which is actually trying to innovate rather than top mere benchmarks. Others like OpenAI, Anthropic and Google are mostly just competeing with each rather than keep innovating around the clock.
They compete with each other by innovating. The innovations result in more utility for the customer, but the technology isn't made public. Trade secrets are secret for a reason.
The reason people may think that DeepSeek is the "most innovative" is because of what they can observe from the outside, much like people may mistakenly conclude models are the "prettiest of the population" because not everyone is photographed for public consumption.