Live data from Hacker News

DSpark: Speculative decoding accelerates LLM inference [pdf]

github.com

101–110 of 393 posts

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#101

DeepSeek continues to not only push the boundaries but also publish these incredible papers explaining how they achieved their gains - something the American labs no longer do unfortunately. Chinese labs are doing the most interesting work in AI right now.

Publishing by necessity I wonder? American labs on the cutting edge pioneering the way forward, so Deepseek open sourcing what they’ve got is to help even the playing field. Hopefully the experts here can offer insight. The above is just my hunch and I’m not a specialist in this field.

> Publishing by necessity

It's more a cultural thing. Sharing progress is just in their blood.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#102
post #97

Earlier quoted context omitted.

It's even worse than that. China publishes stacks upon stacks of policy documents in which they explain clearly what they will do and why. This includes why they do poverty alleviation and why they believe big monopolies that own everything are bad. But almost no western observers care to read those documents. Instead, western observers, including HN, speculate endlessly about China's intentions, and "it would be nai…

I 100% agree with you and want to add something. If you simply take what the Chinese government says at face value, you will be correct way more often than 95% of Western policy wonks, media talking heads, "analysts" and so forth. Because, like you say, they tell you everything they're doing. In the recent US-China summit, Xi Jinping just came out and used the Thucydides Trap metaphor, which tells you everything abou…

The Thucydides Trap mention is different from what you describe. Xi has dismissed the Thucydides Trap multiple times in the past as being hearsay and self-imposed bias (https://www.globaltimes.cn/content/944179.shtml). "We should strictly base our judgment on facts, lest we become victims to hearsay, paranoid or self-imposed bias. There is no such thing as the so-called Thucydides trap in the world. But should major countries time and again make the mistakes of strategic miscalculation, they might create such traps for themselves."

But western politicians keep raising this metaphor. So at some point they're like "okay we'll speak your language". They then used this metaphor to make the point "our rise isn't the threat, your fear of it is. If you resist it you're walking right into the trap Thucydides warned about". So your conclusion is still right, they don't want open hostilities, a stable world is in their interest.

Then western media ran away with this and were like "OMG Xi mentioned the Thucydides Trap", completely ignoring his point.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#103
post #92
post #67

Earlier quoted context omitted.

Only reliable way to have zero data retention is to self-host.

True. But at some point you got to close your eyes and take a step forward. It’s like with VPN providers. Is Mullvad actually collaborating with law enforcement? They very well could be. It is a calculated risk. Is DeepInfra actually logging and training or selling the logs? They could be.

Mullvad has proved it doesn't collect. It's laughable to even suggest it.

They have been raided multiple times, tons of audits, does bleeding edge research on privacy preserving tech, donates to GOS, etc etc. You don't see this kind of VPN company at all because none exists.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#104
post #87
post #85

Earlier quoted context omitted.

Who is financing DeepSeek and what are they expecting in return?

They are self financed, the company that makes DeepSeek is a finance company that trades on the markets.

Even if they were fully self-financed, which isn’t the case, they would expect something in return.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#106

DeepSeek continues to not only push the boundaries but also publish these incredible papers explaining how they achieved their gains - something the American labs no longer do unfortunately. Chinese labs are doing the most interesting work in AI right now.

Sure, in part by "stealing" from American AI companies with Distillation attacks: https://yipzap.com/anthropic-accuses-alibaba-of-largest-ai-d...

If your moat is “please don’t copy my outputs”, you don’t have a moat. There is no such thing as a distillation “attack”.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#107
post #99
post #69

At this point why can't someone produce a fridge or container-sized AI appliance based on legacy chips (12nm)? I imagine this would cover 80% of corporate use cases where you need to "google-in-a-box" functionality. The state-of-the-art nanometer are impossible to achieve but if you have infinite solar energy during business hours does it really matter? Every company has a parking spot so this ASIC-like appliance cou…

See "exabox" from George Hotz: https://tinycorp.myshopify.com/products/exabox-preorder

No one's buying that shitbox.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#108
post #54
post #39

Earlier quoted context omitted.

Very interesting take

It's a standard take since it is how markets tend to work. They aren't powered by altruism, it is a big system for turning greed into good results. We don't have all this stuff because people suddenly woke up one morning and decided to be nice.

Yes but there's more to the world than markets.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#109

DeepSeek continues to not only push the boundaries but also publish these incredible papers explaining how they achieved their gains - something the American labs no longer do unfortunately. Chinese labs are doing the most interesting work in AI right now.

Google and Microsoft publish more than enough and American universities are publishing the science beyond DeepSeek's engineering. That fact that you don't know about them means you're not following the science only reading hacker news.
Post reply on HN