Live data from Hacker News

OpenAI Jalapeño: Better than Nvidia Blackwell

newsletter.semianalysis.com

301–310 of 389 posts

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#301

Continued hardware improvements really make it hard for me to believe token prices will not continue to plummet.

> continue to plummet.

Continue what? The cost per output token has kept going up for the past three years across the board, as thinking models keep leaning more on test-time scaling.

The quality of the said output tokens obviously increased, and arguably increased more than their price, but the price still went up. Or, on the flip side, the price of combined tokens went down (a bit, it did not "plummet" at all though) but so did the average token quality if you count thinking tokens.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#302

These nascent inference chip efforts are reminding me of the early 3dfx / riva / mach / powervr days. Will be interesting to see if inference chips are here to stay and, if so, who the eventual dominant player(s) will be

I’d go with “whoever owns the entire vertically integrated ecosystem” - which is where the straight GPU comparison falls flat.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#303
post #238

Earlier quoted context omitted.

But that means your different chips all have different sets of weights and are different generations. If none of that is baked into the chip as now then all the chips are running the latest weights every time. Even if you could ignore the stuff built into the chip when the time came, at that point you just wasted money on silicon that’s useless in 2-3 months.

Imagine Anthropic gives you Opus of 6 months ago but at much higher speeds and much lower cost (that they might or might not pass on). Would you use it?

Yes, this would basically obviate Sonnet and Haiku. If you consider them 1 and 2 generations behind, respectively (that's not really what they are), you can still get a ton out of those older chips. Not to mention people still use older Opus versions happily. (In part because they don't like the new Opus but still, the cost effectiveness is a huge boon.)

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#304

Earlier quoted context omitted.

I do believe that this is the trade off. We are more efficient but slower in terms of thinking (at the same level of intelligence). Some animals go much further in terms of that trade off, see https://en.wikipedia.org/wiki/Portia_(spider) for example.

Just checked wikipedia page... they have like 100K neurons only.. wtf!!! How can nature cramp all senses, including spatial, motion, life maintenance and general thinking into just 100K neurons?

100k is already a lot. Plenty of insects get by with a couple of thousand, and pack a whole bunch of complex behaviour, including flight, into that.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#305
post #85

I think they talked about this being general purpose chip but I would think that Anthropic/OpenAI are at the scale now they could bake LLM weights into chips themselves. For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough. While 2 years ago nothing was useful more than 1 year long, there are many older m…

"Baking in" a model into a chip is a bad idea because chips take 2 years to tape out and then you're stuck doing inference on llama 3 in 2026 when fable/sol are available. Every accelerator is a tradeoff between flexibility and performance and GPUs are already pareto-optimal

Closer to 2 weeks, as long as the new model fits: your model bits are entirely in mask rom, so you can change them with a metal only ECO that only touches two metal layers. You probably don't even need to run any timing analysis, etc. since all of the changes are going to be isolated in a gigantic square that's isolated from everything else.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#306
post #96

I hadn't seen the token/Joules comparison with human speech before. Humans are still 22x more efficient, which is not that far considering the rate of progress in this area.

The 20W number includes EVERYTHING else the brain does. The chips/models are literally only producing tokens. Let's see an LLM drive a robot harness and have the robot produce speech, as well as move through 3D space, keep track of metabolic needs, etc. etc. etc. before we compare efficiencies. That is even assuming the tokens are of equal quality. This comparison is currently Apples and Oranges.

> At low concurrency scenarios, Jalapeño demonstrates remarkable interactivity, hitting over 700 tokens per sec per user at concurrency 1 on the DeepSeek R1 model.

This is about the same rate you get out of Sol Ultraspeed.

Why do you think extra tool calls like that would be so unthinkable? It'd run circles around this, especially if the problem can be split up among a live-collaborating agent swarm, so that it's not a single user thing anymore, which is exactly what they have in the cooker with Astra.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#308

When people talk about the commodification of inferencing, they imagine a future where everyone has access to frontier models and can run them at the same cost, and what will actually happen is closer to the commodification of _oil_, where only a few companies have the scale to produce it at a competitive price, and advances like this are _why_. Once models are more or less interchangeable, the price of LLMs will dro…

I don't believe models will be commodified because each model is unique with strengths and weaknesses. Its not like Steel which is more or less the same no matter where you purchase it from. If what you said were true, you would hardly see people complaining about the quality of Opus 5 or good writing from Sol. But people do.

Steel has different varieties like carbon steel, alloy steel, stainless steel, and tool steel.

Each of those categories then has different grades of quality.

Tokens are a lot more like steel than oil, especially since a lot of tokens are used as structural material in the form of code.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#309

Earlier quoted context omitted.

Probably! But not viable yet; the chips would be about a year behind SOTA. Note the ~16 months that the article quotes as being insanely fast to get this chip to tape-out (read: start producing). We'll have to bootstrap our way there: AI is actively being used to get us closer to viable lead times for this. Unfortunately, there's some real physical constraints: IIRC, manufacturing a wafer takes on the order of a mont…

I think Sol is already good enough though.

I don't think this is true at all. I use opus 5 to build a small project. It bothers me that people simply handwave away concerns so easily.

The context size is still not nearly as big as it needs to be to store all the code an enterprise needs. And apparently as you increase the context size, there are more defects. So none of this is a solved problem. We are still in early days and there is a lot to be done.

I'm sure there are marketing people who will say "coding is solved" and other such snake oil but none of this is done, far from it!

That being said there is still enormous value in older models especially with tool calling which will let them access the latest data. I feel like we need to be a little more careful and the tools should cache in a smart way to avoid rework but clearly if we could have opus 4.8 level of work for like a one time payment of a system for local LLM it will have value for years into the future.

So I agree in a weird way that sol is good enough for certain tasks but really there is a long road ahead.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#310
post #13

It's so funny to see FP4.... I remember 20 years ago being asked what sort of HPC we needed in genomics, and the answer was basically, "lower precision, faster" for the stuff I was working on. But FP4 is, well, almost comical. One thing not on that comparison table: die size. If I'm understanding that correctly, it's about the same as the Rubin, but at 1/3 the number of NVFP4 PFLOPs. (The text disagrees with the tabl…

Agree, I remember when even half precision made its way into C# sometime around 2020 (I didn’t know much about ML then) and I thought, well I guess that’s a worthwhile tradeoff but I can’t imagine going lower. Lo and behold (1-bit Bonsai) how much lower you could go.

But Bonsai is garbage.
Post reply on HN