Earlier quoted context omitted.
They are self financed, the company that makes DeepSeek is a finance company that trades on the markets.
Even if they were fully self-financed, which isn’t the case, they would expect something in return.
DSpark: Speculative decoding accelerates LLM inference [pdf]
161–170 of 393 posts
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#162Earlier quoted context omitted.
I don’t see how Anthropic is in a better position. They have a slight edge in model quality right at a time when we’re getting a taste of what cheap, “good enough” AI looks like. They don’t own their own compute. And their own arrogance and lies have alienated a huge chunk of their customer base and alerted everyone to the dangers of being dependent on them.
I personally think not owning their own compute is going to be an advantage. There is a meteor headed towards all this AI investment that I don't think has been properly accounted for and that is, what happens to all the existing hardware investments when NVidia's next architecture comes out. Blackwell (H100/H200) is the current generation. Rubin (R100, presumably R200) is the next and arrives soon. Now a lot of the…
However, Google probably won't catch up. Nvidia has been winning in spite of the fact that their hardware is general purpose rather than tuned for inference.
Rubin has architectural differences I don't understand that are supposed to make inference much cheaper and faster while still retaining those other more generic capabilities. Their next generation after that is going to do even better at being fast for inference and general purpose.
Google is betting that their TPUs won't depreciate faster than the markup they have to pay to Nvidia. I don't think they will be right.
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#163Earlier quoted context omitted.
And humans don't run on markets.
Mostly they kind of do since we do live in an utopian society of unlimited abundance. Extremely few people can afford to (or want to) spend a very large number of working hours without ever getting anything directly in return for it.
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#164DeepSeek is, as I feel currently, the sole AI company which is actually trying to innovate rather than top mere benchmarks. Others like OpenAI, Anthropic and Google are mostly just competeing with each rather than keep innovating around the clock.
Qwen as well.
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#165Earlier quoted context omitted.
Chinese labs are also still behind, so they’re incentivized to collaborate and have no reason to do it in private. I suspect their tune will change if they ever take the lead..
The question is also what game they're playing. Deepseek came out of a hedge fund. I think it's no coincidence that their publications tend to have a large impact on AI stock prices. Destroying the growth story of overvalued stocks is an interesting investment strategy. It's not even new. Shortsellers understandably get terrible rep from execs, but their actions are more often in the public interest than you'd think.…
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#166DeepSeek continues to not only push the boundaries but also publish these incredible papers explaining how they achieved their gains - something the American labs no longer do unfortunately. Chinese labs are doing the most interesting work in AI right now.
Exactly. They did not have to open up their research up and this is what happens when smart researchers are forced to squeeze performance gains out of existing hardware. They don't have TPUs or access to the latest Vera Rubin GPUs either to get performance gains for free. All of the optimizations Deepseek have done are in software and it goes down to the PTX assembly level. Compared to Anthropic who are celebrating i…
DeepSeek are still using NVIDIA (PTX) to train on, but for inference have already transitioned to Huawei Ascend chips, and inference speed is what this paper is addressing.
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#167DeepSeek continues to not only push the boundaries but also publish these incredible papers explaining how they achieved their gains - something the American labs no longer do unfortunately. Chinese labs are doing the most interesting work in AI right now.
Exactly. They did not have to open up their research up and this is what happens when smart researchers are forced to squeeze performance gains out of existing hardware. They don't have TPUs or access to the latest Vera Rubin GPUs either to get performance gains for free. All of the optimizations Deepseek have done are in software and it goes down to the PTX assembly level. Compared to Anthropic who are celebrating i…
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#168Earlier quoted context omitted.
Anthropic almost certainly also has optimized software down to the assembly level, considering this take-home interview challenge they published: https://github.com/anthropics/original_performance_takehome/... which is all about instruction-level performance optimizations. That they don't prioritize UI fixes just means they consider other things more important.
that's pretty silly to use as a measure of what they do internally
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#169Earlier quoted context omitted.
Wouldn’t that just help the American labs anyway though? Or do they assume they’ve actually already figured this stuff out and kept it secret?
It used to be the case that NSA hired the majority of all math graduates in the US, and were assumed to be years ahead in cryptography. Yet in the 90s, it became clear that they no longer were that - among other things, the cipher of the notorious Clipper chip was broken, and we can rule out that it was made weak on purpose because the whole point of Clipper was that they had a backdoor. So, despite hiring the cream…
Re: DSpark: Speculative decoding accelerates LLM inference [pdf]
#170Earlier quoted context omitted.
Chinese companies (and labs) operate in conjunction with the CCP so whatever they're doing, it's because it's Chinese state policy. What became clear when DeepSeek came onto the scene was that China was seeking to commoditize LLMs. They consider it an issue of national security not to be beholden to US tech companies when it comes to AI. And I, for one, fully endorse this policy. Another data point on this is the bla…
> Another data point on this is the black market for Claude tokens in China [1]. The chat logs themselves are a commodity to train models. anyone with IQ higher than 130 (thus qualified for actual AI R&D) would be questioning something obvious here - if they are already doing such dodgy stuff with the aim to maximize profits, why would those resellers have large amount of logs with actual American model responses to…