Live data from Hacker News

OpenAI Jalapeño: Better than Nvidia Blackwell

newsletter.semianalysis.com

201–210 of 389 posts

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#202
post #96

Earlier quoted context omitted.

The 20W number includes EVERYTHING else the brain does. The chips/models are literally only producing tokens. Let's see an LLM drive a robot harness and have the robot produce speech, as well as move through 3D space, keep track of metabolic needs, etc. etc. etc. before we compare efficiencies. That is even assuming the tokens are of equal quality. This comparison is currently Apples and Oranges.

Right, I can do the talked about ~3 tok/sec output and drive a car, hold my bladder, and eat chips at the same time. Take that, Jalapeno!

For what it's worth, LLMs don't really suffer from incontinence, so at least that part is pretty much a solved problem.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#203
post #186

I hadn't seen the token/Joules comparison with human speech before. Humans are still 22x more efficient, which is not that far considering the rate of progress in this area.

> Humans are still 22x more efficient, which is not that far considering the rate of progress in this area. Based on a human output rate of 3.3 tok/s, which seems questionable as a means of comparison

What exactly are you questioning?

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#204
post #186

Earlier quoted context omitted.

> Humans are still 22x more efficient, which is not that far considering the rate of progress in this area. Based on a human output rate of 3.3 tok/s, which seems questionable as a means of comparison

What exactly are you questioning?

The claim that tok/s independent of quality is a useful comparison (I can get thousands of tok/s on a suitable small model), and secondarily that humans can’t output “tokens” faster than than in some sense, which I am less confident about

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#205
post #85

I think they talked about this being general purpose chip but I would think that Anthropic/OpenAI are at the scale now they could bake LLM weights into chips themselves. For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough. While 2 years ago nothing was useful more than 1 year long, there are many older m…

"Baking in" a model into a chip is a bad idea because chips take 2 years to tape out and then you're stuck doing inference on llama 3 in 2026 when fable/sol are available. Every accelerator is a tradeoff between flexibility and performance and GPUs are already pareto-optimal

Well we'll see those surplus chips being repurposed for toys then. Who wouldn't want a new Furby that can actually hold a conversation.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#206

Earlier quoted context omitted.

Yes, they could also sell me GPT Sol 5.6 or 5.7 on a chip and I’d probably buy it. It’s a really really useful model for me, I’m not sure how much better for coding I need it to be. For most things I find Sol good enough with a small amount of coaxing around my tastes.

Man wouldn’t it be cool to be able to slot a massive ROM AI chip into the external AI drive of the pc…

It'd just be pcie probably

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#207

Competition is good for all of us, we will get better and faster chips. Or at least Nvidia GPUs will become slightly cheaper for regular consumers again

That's if any datacenters are allowed to be built with them. There is probably a ~50% chance that the next Dem candidate for presidency runs on a national datacenter moratorium or something equally as crippling.

If the populist campaign is to Make Affordable DRAM Again, then it's not a terrible solution.

The current datacenter owners love a compute-bound world anyhow. A moratorium on new datacenters would increase their valuation, encourage efficiency and make computers cheap again. If Chinese labs can ship frontier models under 1T parameters, why not American labs too?

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#208

Earlier quoted context omitted.

Right, I can do the talked about ~3 tok/sec output and drive a car, hold my bladder, and eat chips at the same time. Take that, Jalapeno!

For what it's worth, LLMs don't really suffer from incontinence, so at least that part is pretty much a solved problem.

They sometimes leak their system prompt

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#209

Earlier quoted context omitted.

"Baking in" a model into a chip is a bad idea because chips take 2 years to tape out and then you're stuck doing inference on llama 3 in 2026 when fable/sol are available. Every accelerator is a tradeoff between flexibility and performance and GPUs are already pareto-optimal

It depends when the good enough level hits. Pretty sure we are almost there for most common applications of AI.

That's only half the problem. OpenAI is contractually obligated, if you will, to believe that models will continue improving at an impressive rate for the foreseeable future (otherwise their valuation makes no sense).

If you believe that, then you should expect to get Sol-level performance out of a Luna-cost model within six months or a year. If you have a system with the weights baked in, that means you're going to end up serving that Sol-class model several times more expensively than it will take someone who comes along a few months later. (such as what recently happened with DeepSeek's update.)

And under that assumption of continuing advancement, baking things in doesn't make sense in general - it's a play you'd make if you think things are slowing down a lot. Which may be right but it's not OpenAI or anthropic's play.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#210
post #186

I hadn't seen the token/Joules comparison with human speech before. Humans are still 22x more efficient, which is not that far considering the rate of progress in this area.

> Humans are still 22x more efficient, which is not that far considering the rate of progress in this area. Based on a human output rate of 3.3 tok/s, which seems questionable as a means of comparison

I do believe that this is the trade off. We are more efficient but slower in terms of thinking (at the same level of intelligence). Some animals go much further in terms of that trade off, see https://en.wikipedia.org/wiki/Portia_(spider) for example.
Post reply on HN