Live data from Hacker News

OpenAI Jalapeño: Better than Nvidia Blackwell

newsletter.semianalysis.com

121–130 of 389 posts

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#121

Competition is good for all of us, we will get better and faster chips. Or at least Nvidia GPUs will become slightly cheaper for regular consumers again

If you think tanking Trillions in investments, warming the earth and increasion ocean water levels, creating water shortages and brown-outs is "good for all of us" - well, the rest of us beg to differ.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#123

These nascent inference chip efforts are reminding me of the early 3dfx / riva / mach / powervr days. Will be interesting to see if inference chips are here to stay and, if so, who the eventual dominant player(s) will be

Which in turn reminds me of Soundblaster audio cards! I suspect inference chips are closer to the GPU story than the Soundblaster story though.

I remember one soundblaster card I bought came with a Lara Croft demo, that exploited the incredible immersion of real time dynamic reverb.

Genuinely I think game audio took a few steps back from that heady era, the innovation in audio likely didn't sell as many cards as graphics innovations did.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#124
post #96

Earlier quoted context omitted.

The 20W number includes EVERYTHING else the brain does. The chips/models are literally only producing tokens. Let's see an LLM drive a robot harness and have the robot produce speech, as well as move through 3D space, keep track of metabolic needs, etc. etc. etc. before we compare efficiencies. That is even assuming the tokens are of equal quality. This comparison is currently Apples and Oranges.

And the brain is literally only producing electrochemical signals. I don’t see how tokens can’t produce speech or track metabolic needs. You can talk to chatgpt can’t you? Or do you mean literally talking? Because that’s not a brain function, that’s the mouth, vocal chords, and lungs.

> I don’t see how tokens can’t produce speech or track metabolic needs.

It probably could, but the point is this would require additional tokens, blowing up the comparison. The token output of LLMs and "token output" of speech are simply at different abstraction levels. Hence my comparison to the LLM brain driving the robot harness to produce speech etc. This would be more comparable, and also look significantly worse than "only" the 22x less efficient number.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#125

Competition is good for all of us, we will get better and faster chips. Or at least Nvidia GPUs will become slightly cheaper for regular consumers again

If you think tanking Trillions in investments, warming the earth and increasion ocean water levels, creating water shortages and brown-outs is "good for all of us" - well, the rest of us beg to differ.

these GPUs make computation faster, I understand as of now maybe all the computation is used to generate yet another junk LinkedIn post or unnecessary RFC, but at some point this craze should settle and we will be left with powerful computation machines, which can be used for computing more useful things

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#126
post #31

Continued hardware improvements really make it hard for me to believe token prices will not continue to plummet.

This may just be a classic case of Jevons paradox: https://en.wikipedia.org/wiki/Jevons_paradox In short, better hardware will drive down token cost in the near-term, but will drive up the demand for tokens as it gets cheap enough for other sectors to start to use it heavily. It comes from steam engines where economists originally thought that coal demand would plummet with more efficient engines, but it actually jus…

That's when demand is higher than capacity. Now imagine places like Gigalab and Chinese labs are online and able to produce significant percentage of chips. That could cause real surge in prices.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#127

I hadn't seen the token/Joules comparison with human speech before. Humans are still 22x more efficient, which is not that far considering the rate of progress in this area.

I couldn't source the parameters from the screenshot or the nearby graphs, but from the nearby graphs you can see that at concurrency C=1, tokens/Joule (vertical axis) has totally plummeted, and obviously concurrent inference is much more efficient by batching. Divide the memory by the bandwidth and thats how long it takes to dump the full RAM contents through the chip. Do you want to do this once per token for a single conversation, or do you want to progress multiple conversations if you're going through all the weights anyway? The peak in the graphs is easily 22x more efficient than the low bottom right part on the graphs. So in batched mode its already more efficient than human speech.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#128
post #13

Earlier quoted context omitted.

Agree, I remember when even half precision made its way into C# sometime around 2020 (I didn’t know much about ML then) and I thought, well I guess that’s a worthwhile tradeoff but I can’t imagine going lower. Lo and behold (1-bit Bonsai) how much lower you could go.

Ternary?

Knuth's base-e proposal enters the chat.

They were right about everything 50+ years ago, but they didn't have the budget for the right hardware, had to write conference papers and books instead.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#129

When people talk about the commodification of inferencing, they imagine a future where everyone has access to frontier models and can run them at the same cost, and what will actually happen is closer to the commodification of _oil_, where only a few companies have the scale to produce it at a competitive price, and advances like this are _why_. Once models are more or less interchangeable, the price of LLMs will dro…

I don't agree. At the moment companies like NVIDIA take several times what it costs to make a chip. I think the fair split for the technology contribution is more like 50-50, maybe even 30-70 in favour of the manufacturer. With competition we will actually have the fair split, whatever that is, and thus much lower prices. At the moment, to have a big AI firm, or really AI firm at all, you need to be blessed by NVIDIA…

10-90 is the fair split.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#130
post #123

These nascent inference chip efforts are reminding me of the early 3dfx / riva / mach / powervr days. Will be interesting to see if inference chips are here to stay and, if so, who the eventual dominant player(s) will be

Which in turn reminds me of Soundblaster audio cards! I suspect inference chips are closer to the GPU story than the Soundblaster story though. I remember one soundblaster card I bought came with a Lara Croft demo, that exploited the incredible immersion of real time dynamic reverb. Genuinely I think game audio took a few steps back from that heady era, the innovation in audio likely didn't sell as many cards as grap…

On board got "good enough" and the separate cards died away.

In fairness on board (depending on the board but on the whole) is pretty good.

Post reply on HN