Live data from Hacker News

DSpark: Speculative decoding accelerates LLM inference [pdf]

github.com

81–90 of 393 posts

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#81
post #39

Earlier quoted context omitted.

Very interesting take

I don't understand what is interesting about it: it's the default. Markets don't run on altruism.

The standard is applied very inconsistently. Nobody accuses the local bakery of being motivated by profit, and that they don't bake bread for you out of altruism.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#82
post #66
post #50

Earlier quoted context omitted.

Chinese companies (and labs) operate in conjunction with the CCP so whatever they're doing, it's because it's Chinese state policy. What became clear when DeepSeek came onto the scene was that China was seeking to commoditize LLMs. They consider it an issue of national security not to be beholden to US tech companies when it comes to AI. And I, for one, fully endorse this policy. Another data point on this is the bla…

> Another data point on this is the black market for Claude tokens in China [1]. The chat logs themselves are a commodity to train models. anyone with IQ higher than 130 (thus qualified for actual AI R&D) would be questioning something obvious here - if they are already doing such dodgy stuff with the aim to maximize profits, why would those resellers have large amount of logs with actual American model responses to…

The galaxy brains in the labs putatively buying the logs wouldn't notice this? Or figure out a structure to prevent this?

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#83
post #30

Earlier quoted context omitted.

Which is a good thing. Self-serving motives are more reliable than altruistic ones.

The world runs on incentives. Altruism/Self-serving are down stream of that. Wikipedia is altruistic, and serves humanity quite well.

Open-source is also altruistic. If DeepSeek does become self-serving once they get the top spot, it doesn’t take away from the altruistic contributions that they made towards open models.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#84
post #63

Earlier quoted context omitted.

Chinese labs are also still behind, so they’re incentivized to collaborate and have no reason to do it in private. I suspect their tune will change if they ever take the lead..

Projection is a funny thing. It causes people to misread situations all the time. Southern slaveowners feared violent retribution from freed slaves, for example [1]. It was pure projection and said more about the South than it did the slaves. The reality was there was no violent retribution. It was the opposite where the former slaveowners continued to inflict violence on the formerly enslaved. I say this because we…

It's even worse than that. China publishes stacks upon stacks of policy documents in which they explain clearly what they will do and why. This includes why they do poverty alleviation and why they believe big monopolies that own everything are bad. But almost no western observers care to read those documents. Instead, western observers, including HN, speculate endlessly about China's intentions, and "it would be naive to believe they would not do X" or drawing equivalences to Soviet Union or whatever. And the "journalists" sell this notion that Chinese state intentions are "untransparent" and "unknowable" while pretending the policy documents don't exist.

Meanwhile, Xi Jinping has published his 5th book on how governance in China works and what they're after. These are not books written for a western audience: they're compilations of speeches that he already gave to the Chinese party and state apparatus, so the contents are not sanitized for foreign audiences. But there are no English reviews of summaries of this 5th book at all by the usual China experts that distribute what western audience know about China.

This extends to beyond the government. Even though "for the people but only against the government" is an often-heard mantra, nobody seems to listen to what Chinese AI companies themselves say about why they publish open models. DeepSeek and GLM have said multiple times publicly what their motivations are, yet people on HN still speculate like they usually do.

Truly mind-boggling. I get that a lot of people don't like China. But setting aside the question of whether their dislike is justified, it would at least be rational to properly understand China, even if it's to defeat it. And listening to what China says themselves is absolutely essential for proper understanding. But people don't bother to? And they seem mostly happy with sticking to speculations that match preconceived notions, even if that hurts their chances of defeating China.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#85

DeepSeek continues to not only push the boundaries but also publish these incredible papers explaining how they achieved their gains - something the American labs no longer do unfortunately. Chinese labs are doing the most interesting work in AI right now.

Probably because American AI companies are on the hook for quite a lot of investment money. I think they are trying to find the magical moat to justify their valuation. Revealing optimizations similar to these would pretty much reduce their competitive position.

Who is financing DeepSeek and what are they expecting in return?

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#86

DeepSeek continues to not only push the boundaries but also publish these incredible papers explaining how they achieved their gains - something the American labs no longer do unfortunately. Chinese labs are doing the most interesting work in AI right now.

R1 was very influential on US models development.

[deleted]

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#87
post #85

Earlier quoted context omitted.

Probably because American AI companies are on the hook for quite a lot of investment money. I think they are trying to find the magical moat to justify their valuation. Revealing optimizations similar to these would pretty much reduce their competitive position.

Who is financing DeepSeek and what are they expecting in return?

They are self financed, the company that makes DeepSeek is a finance company that trades on the markets.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#88
post #50

DeepSeek continues to not only push the boundaries but also publish these incredible papers explaining how they achieved their gains - something the American labs no longer do unfortunately. Chinese labs are doing the most interesting work in AI right now.

Chinese companies (and labs) operate in conjunction with the CCP so whatever they're doing, it's because it's Chinese state policy. What became clear when DeepSeek came onto the scene was that China was seeking to commoditize LLMs. They consider it an issue of national security not to be beholden to US tech companies when it comes to AI. And I, for one, fully endorse this policy. Another data point on this is the bla…

I don’t see how Anthropic is in a better position. They have a slight edge in model quality right at a time when we’re getting a taste of what cheap, “good enough” AI looks like. They don’t own their own compute. And their own arrogance and lies have alienated a huge chunk of their customer base and alerted everyone to the dangers of being dependent on them.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#89

Earlier quoted context omitted.

Publishing by necessity I wonder? American labs on the cutting edge pioneering the way forward, so Deepseek open sourcing what they’ve got is to help even the playing field. Hopefully the experts here can offer insight. The above is just my hunch and I’m not a specialist in this field.

Wouldn’t that just help the American labs anyway though? Or do they assume they’ve actually already figured this stuff out and kept it secret?

It used to be the case that NSA hired the majority of all math graduates in the US, and were assumed to be years ahead in cryptography. Yet in the 90s, it became clear that they no longer were that - among other things, the cipher of the notorious Clipper chip was broken, and we can rule out that it was made weak on purpose because the whole point of Clipper was that they had a backdoor.

So, despite hiring the cream of the crop of math graduates, who could read the papers of free academia, but whose own result the free world could not access - they fell behind.

I have a theory explaining why. I think it's because science is an interactive process. NSA cryptographers could read papers, but they couldn't talk openly with the authors of those papers, because of secrecy demands - even asking question might indicate what they were working on. You can easily imagine them spending months on something they could have avoided by going to the original authors and getting told "Oh, we tried that for a long time, it doesn't work".

Whether that theory is right or not, cryptography is a concrete example of a domain where public research with fewer resources beat private research with a lot more resources.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#90

DeepSeek continues to not only push the boundaries but also publish these incredible papers explaining how they achieved their gains - something the American labs no longer do unfortunately. Chinese labs are doing the most interesting work in AI right now.

Publishing by necessity I wonder? American labs on the cutting edge pioneering the way forward, so Deepseek open sourcing what they’ve got is to help even the playing field. Hopefully the experts here can offer insight. The above is just my hunch and I’m not a specialist in this field.

Yes, challenger Labs publish out of necessity. It is a marketing strategy. People assuming open source means giving something up, but the reality is that Z.ai has a revenue of some $100M and it would be about $0M if they never open sourced their models.
Post reply on HN