Live data from Hacker News

OpenAI Jalapeño: Better than Nvidia Blackwell

newsletter.semianalysis.com

351–360 of 390 posts

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#351
post #96

Earlier quoted context omitted.

The 20W number includes EVERYTHING else the brain does. The chips/models are literally only producing tokens. Let's see an LLM drive a robot harness and have the robot produce speech, as well as move through 3D space, keep track of metabolic needs, etc. etc. etc. before we compare efficiencies. That is even assuming the tokens are of equal quality. This comparison is currently Apples and Oranges.

> At low concurrency scenarios, Jalapeño demonstrates remarkable interactivity, hitting over 700 tokens per sec per user at concurrency 1 on the DeepSeek R1 model. This is about the same rate you get out of Sol Ultraspeed. Why do you think extra tool calls like that would be so unthinkable? It'd run circles around this, especially if the problem can be split up among a live-collaborating agent swarm, so that it's not…

They are not. If the robot speech is a tool call, then for a fair comparison we need to take the tool call scaffolding (and probably the reasoning too) into account. So rather than a sentence of 10 tokens worth of speech being the output, the raw token output would be maybe 10x or 100x that. Even more if we consider the management of other aspects of the robot embodiment (or we reduce the brain's 20W number to whatever is actually required to produce coherent speech, sadly it is all rather entangled so this is not so easy).

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#352
post #348

Earlier quoted context omitted.

Doesn't matter if the chip is 100x-1000x more efficient and faster, and you can just make a new one for new weights. Imagine being able to run GPT Sol at 1k tokens/second a year from now, at a 100x lower cost per token than now. Would that be useful? Or Qwen 3.8 27B at 10k tokens/second. The super long thinking that makes qwen so effective would take a couple of seconds.

> Imagine being able to run GPT Sol at 1k tokens/second a year from now, at a 100x lower cost per token than now. Would that be useful? Given the current rate of change, it would be hard to guess either way. By some measures the cost at fixed quality score goes down vastly faster than that: A similar trend is evident in the cost of models scoring above 50% on GPQA, a substantially more challenging benchmark than MMLU…

This is from 2025, how about the last 6 months?

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#353
post #85

I think they talked about this being general purpose chip but I would think that Anthropic/OpenAI are at the scale now they could bake LLM weights into chips themselves. For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough. While 2 years ago nothing was useful more than 1 year long, there are many older m…

Eventually, someone is going to do this in Minecraft

Lol i'm surprised this hasn't happened already. Minecraft is a great platform for scripting as long as you think every single other platform is too fast and convenient.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#354

Earlier quoted context omitted.

It may not matter. Think about why SOTA model companies are exploring chips. What do chips offer? If SOTA models haven’t peaked, then the SOTA model companies would still be churning out better and better intelligence.

Google rolled out TPUs in 2015. AWS released Inferentia and Trainium chips in 2020. If companies working on ML-specific chips was evidence that large transformer models have fully saturated their potential, the field would have been done circa GPT-2.

Neither of those companies core business model was serving llms

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#355
post #348

Earlier quoted context omitted.

> Imagine being able to run GPT Sol at 1k tokens/second a year from now, at a 100x lower cost per token than now. Would that be useful? Given the current rate of change, it would be hard to guess either way. By some measures the cost at fixed quality score goes down vastly faster than that: A similar trend is evident in the cost of models scoring above 50% on GPQA, a substantially more challenging benchmark than MMLU…

This is from 2025, how about the last 6 months?

You tell me. Most of these reports take that long to get published, or even longer. Sometimes I even see new-ish reports talking about 4o.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#356
post #161

Earlier quoted context omitted.

Probably! But not viable yet; the chips would be about a year behind SOTA. Note the ~16 months that the article quotes as being insanely fast to get this chip to tape-out (read: start producing). We'll have to bootstrap our way there: AI is actively being used to get us closer to viable lead times for this. Unfortunately, there's some real physical constraints: IIRC, manufacturing a wafer takes on the order of a mont…

The metal masked ROM is basically only 2 metal/contact layers. It's not a full new design and tapeout. You could roll a new set of parameters every ~2-3months. It's not an architectural change. See statements below. https://www.eetimes.com/taalas-specializes-to-extremes-for-e... https://www.turingpost.com/p/taalas https://cambrian-ai.com/taalas-launches-hardcore-chip-with-i... Part of the key is that by moving even f…

[deleted]

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#357
post #252

Earlier quoted context omitted.

Probably! But not viable yet; the chips would be about a year behind SOTA. Note the ~16 months that the article quotes as being insanely fast to get this chip to tape-out (read: start producing). We'll have to bootstrap our way there: AI is actively being used to get us closer to viable lead times for this. Unfortunately, there's some real physical constraints: IIRC, manufacturing a wafer takes on the order of a mont…

They could etch the model architecture, without the weights into the chip. This way newly post-trained model can be loaded and served the same day.

[deleted]

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#358

This semi-analysis article reads a lot more like an OpenAI press release than a real analysis. And to be honest some of the statements seem like just straight up lies - they initially claim they were invited to benchmark it, and then half way down switch to claiming that OpenAI provided all the numbers. This really kind of sucks, because I want to read actual detailed nuanced and credible analysis of what's happening…

Semianalysis is an AI hype organisation, not a serious, unbiased semiconductor reviewer/journalist like chipsandcheese nor a documentarian of the semiconductor industry like Asianonmetry (as it relates so strongly to the modern economies of Asia). If you see something from semianalysis, you can simply ignore it.

I tend to agree, but the Semianalysis + Dwarkesh side of reporting still does surface interesting information. You just have to take it all with a grain of salt, as indeed it is more hype focused, and look for real information hidden in the noise. And possibly to be a bit more entertained as you do so.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#359
post #252

Earlier quoted context omitted.

Probably! But not viable yet; the chips would be about a year behind SOTA. Note the ~16 months that the article quotes as being insanely fast to get this chip to tape-out (read: start producing). We'll have to bootstrap our way there: AI is actively being used to get us closer to viable lead times for this. Unfortunately, there's some real physical constraints: IIRC, manufacturing a wafer takes on the order of a mont…

They could etch the model architecture, without the weights into the chip. This way newly post-trained model can be loaded and served the same day.

[deleted]

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#360

Earlier quoted context omitted.

Semianalysis is an AI hype organisation, not a serious, unbiased semiconductor reviewer/journalist like chipsandcheese nor a documentarian of the semiconductor industry like Asianonmetry (as it relates so strongly to the modern economies of Asia). If you see something from semianalysis, you can simply ignore it.

I tend to agree, but the Semianalysis + Dwarkesh side of reporting still does surface interesting information. You just have to take it all with a grain of salt, as indeed it is more hype focused, and look for real information hidden in the noise. And possibly to be a bit more entertained as you do so.

I suppose access journalism does have to get something out of the bargain, however small it's not zero. If that's the way you like to spend idle time, well it takes all kinds. I don't understand competitive scrabble either.
Post reply on HN