Live data from Hacker News

DSpark: Speculative decoding accelerates LLM inference [pdf]

github.com

291–300 of 393 posts

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#291

Must be wonderful to be on the board of OpenAi et al & their PE investors whilst China keeps blowing up these mines under their feet lmao. Luckily Korean pension funds will buy all the trash as usual but goddamn you gotta start moving quick or you are gonna need some serious AGI to show you how to offload those bonds

Why do you think they have started accusing Chinese labs of stealing and distillation?

A&O no longer have the most to justify their high valuation. The only thing they can do now is to get the government forbid the Chinese models.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#292

Earlier quoted context omitted.

Open-source is also altruistic. If DeepSeek does become self-serving once they get the top spot, it doesn’t take away from the altruistic contributions that they made towards open models.

No parent is right. The core root driver of the world is capitalism, open source exists downstream of that. Software engineers need money to survive. If they exclusively work on open source stuff where are they getting money from to survive? Follow the money trail… even a donation… eventually it leads to an incentive based source or action.

> If they exclusively work on open source stuff where are they getting money from to survive?

These are orthogonal. One can have a paid job while contributing to open source for entirely altruistic reasons.

> Follow the money trail… even a donation… eventually it leads to an incentive based source or action.

BS. Humans do things for altruistic reasons devoid of individual reward all the time.

I, myself, maintain multiple OSS projects entirely for fun and with the hope that others will find it useful. That's it, that's all. I also donate entirely anonymously to charities simply because I believe others deserve support and dignity.

This form of cold, American libertarianism you espouse is pure poison in the body politic, both in this US and globally. It degrades all of human interaction to transactions. Its no wonder that the US is where sociopaths like Zuck were birthed.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#293

Earlier quoted context omitted.

No parent is right. The core root driver of the world is capitalism, open source exists downstream of that. Software engineers need money to survive. If they exclusively work on open source stuff where are they getting money from to survive? Follow the money trail… even a donation… eventually it leads to an incentive based source or action.

> If they exclusively work on open source stuff where are they getting money from to survive? From open source. You can earn money from open source. Open source is not opposed to capitalism, idk where you got that idea.

Young blood, allow me to explain.

I said open source is derivative to capitalism. Meaning open source cannot exist without capitalism. I never said they oppose each other.

Second I said you need to follow the money trail. Money given to people who work on open source comes from non-open source places.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#294
post #245

Earlier quoted context omitted.

Why do people who don't follow the prices of A100 talk like they know things about GPU pricing dynamics? A100s are ~7 years old and going for more than 2 dollars an hour, significantly more expensive than even 2 years ago. This is because anything with 80gb of VRAM or more and made by Nvidia will have economically useful lifespans of like, 10 years. I could see H100s getting 12 years. Micheal Berry doesn't know shit…

So I was curious about how A100s would do running DeepSeek v4. I can't find any instances of running v4 Pro on even an 8xA100 cluster. So you need to run Flash at ~284B params. A100s don't support FP8 so you're running FP16 so you're taking a hit that way. But I see estimates of 30-50tok/s for an 8xA100 cluster. They're drawing 300-400W each so you're looking at probably 3500+ Watts, which is roughly 0.01tok/W. Now j…

You're focussing on inference ... is it not more likely that A100's are being used for training/fine tuning?

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#295

Earlier quoted context omitted.

No parent is right. The core root driver of the world is capitalism, open source exists downstream of that. Software engineers need money to survive. If they exclusively work on open source stuff where are they getting money from to survive? Follow the money trail… even a donation… eventually it leads to an incentive based source or action.

> If they exclusively work on open source stuff where are they getting money from to survive? These are orthogonal. One can have a paid job while contributing to open source for entirely altruistic reasons. > Follow the money trail… even a donation… eventually it leads to an incentive based source or action. BS. Humans do things for altruistic reasons devoid of individual reward all the time. I, myself, maintain mult…

[flagged]

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#296

Earlier quoted context omitted.

> If they exclusively work on open source stuff where are they getting money from to survive? These are orthogonal. One can have a paid job while contributing to open source for entirely altruistic reasons. > Follow the money trail… even a donation… eventually it leads to an incentive based source or action. BS. Humans do things for altruistic reasons devoid of individual reward all the time. I, myself, maintain mult…

[flagged]

>Talking like this is not only against the rules here but it is some of the most vile and direct insults I’ve ever fucking read. And it doesn’t even stem from us disagreeing. It stems from you misunderstanding what was said. Why don’t you read over what I wrote and my explanation before making such a stupid comment. Let me be clear. You’re not stupid, but your reply is stupid. And your reply is stupid because of a misunderstanding. So make yourself understand and clean up your act because shit like this dies worse damage for the world than psychopaths. More wars are started over misunderstanding and uncalibrated anger than actual psychopathic tendencies.

I suggest you practice a much greater degree of self awareness.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#297

Presumably this has been in production for a while, and is one of the reasons they were able to dramatically lower prices a month ago?

good catch, they reduced the prices 75% seems like exactly in line with the speed/inference optimizations gains?

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#298
post #9

This is just one of many papers DeepSeek have released to be able to serve models at extremely cheap prices, unlike the others taking on >$100B+ of debt in building data centers for the same thing. > As with V4-Flash, we treat this point as an indication that DSpark sustains useful throughput under an interactivity target that the baseline cannot efficiently support. At matched system capacities, DSpark delivers 57%…

...... are you really suggesting OpenAI and Anthropic don't have access to these techniques?

if they didn't, they do now. as deepseek published the howto

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#299

Earlier quoted context omitted.

[flagged]

>Talking like this is not only against the rules here but it is some of the most vile and direct insults I’ve ever fucking read. And it doesn’t even stem from us disagreeing. It stems from you misunderstanding what was said. Why don’t you read over what I wrote and my explanation before making such a stupid comment. Let me be clear. You’re not stupid, but your reply is stupid. And your reply is stupid because of a mi…

I am self aware. I’m fully aware of the potential feelings that what I said could evoke but reality is reality and a forum is one of the few places we can still talk about reality without cancel culture or feelings muddling everything up.

I’d rather speak the truth and what I believe in rather than cater to the feelings of people who cannot face objective reality.

And I didn’t openly or directly insult anyone. I criticize where it’s deserved and where it is true. He personally attacked me and I criticized his attack as appropriately as I could.

If you disagree with my premise, attack my argument. Don’t make it personal by telling me to be self aware and replying to a section of my comment not meant for you.

Re: DSpark: Speculative decoding accelerates LLM inference [pdf]

#300

Earlier quoted context omitted.

By this population-only logic, you should concede that India will overtake China. Why not talk about how China shut out American companies for decades before complaining about BYD? As an Indian immigrant, the PRC China has engaged in conflict with almost all its neighbors and stated wars in its short history. China is not so benevolent when they get to the #1 spot: https://m.economictimes.com/industry/renewables/chin…

Its not population only logic, but it does underscore that it is silly to expect the United States to inevitably be ahead. As for the rest of it: https://youtu.be/74DAI2hr9Kk?t=159

I dont necessarily disagree with you on your other points, but the population argument sort of shows a lack of understanding of how China works. It’s not guaranteed at all that they will ever overtake the US.
Post reply on HN