Earlier quoted context omitted.
The 20W number includes EVERYTHING else the brain does. The chips/models are literally only producing tokens. Let's see an LLM drive a robot harness and have the robot produce speech, as well as move through 3D space, keep track of metabolic needs, etc. etc. etc. before we compare efficiencies. That is even assuming the tokens are of equal quality. This comparison is currently Apples and Oranges.
> At low concurrency scenarios, Jalapeño demonstrates remarkable interactivity, hitting over 700 tokens per sec per user at concurrency 1 on the DeepSeek R1 model. This is about the same rate you get out of Sol Ultraspeed. Why do you think extra tool calls like that would be so unthinkable? It'd run circles around this, especially if the problem can be split up among a live-collaborating agent swarm, so that it's not…
OpenAI Jalapeño: Better than Nvidia Blackwell
351–360 of 390 posts
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#352Earlier quoted context omitted.
Doesn't matter if the chip is 100x-1000x more efficient and faster, and you can just make a new one for new weights. Imagine being able to run GPT Sol at 1k tokens/second a year from now, at a 100x lower cost per token than now. Would that be useful? Or Qwen 3.8 27B at 10k tokens/second. The super long thinking that makes qwen so effective would take a couple of seconds.
> Imagine being able to run GPT Sol at 1k tokens/second a year from now, at a 100x lower cost per token than now. Would that be useful? Given the current rate of change, it would be hard to guess either way. By some measures the cost at fixed quality score goes down vastly faster than that: A similar trend is evident in the cost of models scoring above 50% on GPQA, a substantially more challenging benchmark than MMLU…
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#353I think they talked about this being general purpose chip but I would think that Anthropic/OpenAI are at the scale now they could bake LLM weights into chips themselves. For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough. While 2 years ago nothing was useful more than 1 year long, there are many older m…
Eventually, someone is going to do this in Minecraft
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#354Earlier quoted context omitted.
It may not matter. Think about why SOTA model companies are exploring chips. What do chips offer? If SOTA models haven’t peaked, then the SOTA model companies would still be churning out better and better intelligence.
Google rolled out TPUs in 2015. AWS released Inferentia and Trainium chips in 2020. If companies working on ML-specific chips was evidence that large transformer models have fully saturated their potential, the field would have been done circa GPT-2.
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#355Earlier quoted context omitted.
> Imagine being able to run GPT Sol at 1k tokens/second a year from now, at a 100x lower cost per token than now. Would that be useful? Given the current rate of change, it would be hard to guess either way. By some measures the cost at fixed quality score goes down vastly faster than that: A similar trend is evident in the cost of models scoring above 50% on GPQA, a substantially more challenging benchmark than MMLU…
This is from 2025, how about the last 6 months?
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#356Earlier quoted context omitted.
Probably! But not viable yet; the chips would be about a year behind SOTA. Note the ~16 months that the article quotes as being insanely fast to get this chip to tape-out (read: start producing). We'll have to bootstrap our way there: AI is actively being used to get us closer to viable lead times for this. Unfortunately, there's some real physical constraints: IIRC, manufacturing a wafer takes on the order of a mont…
The metal masked ROM is basically only 2 metal/contact layers. It's not a full new design and tapeout. You could roll a new set of parameters every ~2-3months. It's not an architectural change. See statements below. https://www.eetimes.com/taalas-specializes-to-extremes-for-e... https://www.turingpost.com/p/taalas https://cambrian-ai.com/taalas-launches-hardcore-chip-with-i... Part of the key is that by moving even f…
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#357Earlier quoted context omitted.
Probably! But not viable yet; the chips would be about a year behind SOTA. Note the ~16 months that the article quotes as being insanely fast to get this chip to tape-out (read: start producing). We'll have to bootstrap our way there: AI is actively being used to get us closer to viable lead times for this. Unfortunately, there's some real physical constraints: IIRC, manufacturing a wafer takes on the order of a mont…
They could etch the model architecture, without the weights into the chip. This way newly post-trained model can be loaded and served the same day.
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#358This semi-analysis article reads a lot more like an OpenAI press release than a real analysis. And to be honest some of the statements seem like just straight up lies - they initially claim they were invited to benchmark it, and then half way down switch to claiming that OpenAI provided all the numbers. This really kind of sucks, because I want to read actual detailed nuanced and credible analysis of what's happening…
Semianalysis is an AI hype organisation, not a serious, unbiased semiconductor reviewer/journalist like chipsandcheese nor a documentarian of the semiconductor industry like Asianonmetry (as it relates so strongly to the modern economies of Asia). If you see something from semianalysis, you can simply ignore it.
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#359Earlier quoted context omitted.
Probably! But not viable yet; the chips would be about a year behind SOTA. Note the ~16 months that the article quotes as being insanely fast to get this chip to tape-out (read: start producing). We'll have to bootstrap our way there: AI is actively being used to get us closer to viable lead times for this. Unfortunately, there's some real physical constraints: IIRC, manufacturing a wafer takes on the order of a mont…
They could etch the model architecture, without the weights into the chip. This way newly post-trained model can be loaded and served the same day.
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#360Earlier quoted context omitted.
Semianalysis is an AI hype organisation, not a serious, unbiased semiconductor reviewer/journalist like chipsandcheese nor a documentarian of the semiconductor industry like Asianonmetry (as it relates so strongly to the modern economies of Asia). If you see something from semianalysis, you can simply ignore it.
I tend to agree, but the Semianalysis + Dwarkesh side of reporting still does surface interesting information. You just have to take it all with a grain of salt, as indeed it is more hype focused, and look for real information hidden in the noise. And possibly to be a bit more entertained as you do so.