Live data from Hacker News

OpenAI Jalapeño: Better than Nvidia Blackwell

newsletter.semianalysis.com

381–390 of 390 posts

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#381

Earlier quoted context omitted.

They are not. If the robot speech is a tool call, then for a fair comparison we need to take the tool call scaffolding (and probably the reasoning too) into account. So rather than a sentence of 10 tokens worth of speech being the output, the raw token output would be maybe 10x or 100x that. Even more if we consider the management of other aspects of the robot embodiment (or we reduce the brain's 20W number to whatev…

But there are already voice models that do a reasonable job at a fraction of the throughput available? The real question is how expensive it is to coordinate between these different modalities, and I really don't see why it'd be all that much. I half expect Boston Dynamics to show something like this off in Q4 or whatever.

I am not arguing that there are perhaps other models that can run at the same quality, can coordinate between the different modalities, but are way less power hungry. My point is exactly about the comparison between the token output of the LLM running on the jalapeno chip, and sneaking in the power "usage" of the brain in the "token output" of human speech.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#382
post #277

Earlier quoted context omitted.

Since the current NVL72 are still at the ~15%/yr failure rate it's not clear your new data center is going to have half it's compute in 3 years. If you're still running H100s they draw >10x kWh/Mtoken as new designs. All of these systems become dated, but not all of them require entirely new infrastructure. If a ROM rack running a near frontier agent model at >10ktoken/sec costs What these don't do is TRAINING, they…

> Since the current NVL72 are still at the ~15%/yr failure rate it's not clear your new data center is going to have half it's compute in 3 years This would be a stupidly bad failure rate, basically the worst business decision you could make, especially if you're somehow on the hook for eating those losses (which seems to be the implication?). Is there a linkable source on this? The only thing I could find is SemiAna…

I'll note CoreWeave didn't commission the very first Blackwells until ~Jan 2025 so it hasn't been very long for RMAs. Furthermore, for training even when the NVL backplane can use the remaining GPUs a ~20% drop in performance means it isn't in the training cluster.

I've personally heard this from several sources in the data centers (installers, training, network). It's not uncommon for 10% of racks to fail on delivery. I hear that's improved somewhat from GB200 to GB300, but the number of FW updates from the time they ship, until they're commissioned is >>10. If an HBM or GPU or backplane supply/cooling fails, it is basically not swappable or repairable. You have a "dead" rack, and deliveries are on allocation so you don't get a replacement for months (eg some "RMAs" for early delivered parts in late 2025 are still dead racks 9 months later). "Tray" swaps are technically possible, but still quite rare, perhaps because debugging takes as much time as commissioning a new rack.

I don't want to out anyone, but these are similar comments:

https://www.linkedin.com/posts/neelmaster1_aiinfrastructure-...

https://www.hostzealot.com/blog/news/nvidia-gb200-nvl72-is-n...

https://introl.com/blog/gb200-nvl72-deployment-72-gpu-liquid...

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#384

Earlier quoted context omitted.

This would have been far more effective with 1/10 as many words.

Not really. It is a long story.

I really really disagree. Most of the prose is devoid of information. I find it hard to believe the author even read the whole thing one time.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#385
post #347

Earlier quoted context omitted.

I've been reading them since before all of the AI hype, and I've always thought they're pretty good. You a few spicy takes with the overview/opinions/benchmarks. Better than semiaccurate. The article you link says not a lot of criticisms with very many words, and the AI prose gets much worse towards the end, seemingly when the author also gave up on reading it. I am disappointing in the plagiarism though, especially…

It isn't about bunk takes or not. It is about the motivation behind doing something. Their takes are fabricated in such a way as to drive clicks to their business, where they are printing money selling MNDA to the highest bidder. Dylan uses his influence as a service and it is borderline criminal. He just sued a whistleblower employee. It is so blatant, he even lives and works directly with people in power who feed h…

I dunno, it's a blog, so I'm not so worried about the motivation behind it besides how it biases their takes. I think being close to people that feed you information might be prerequisite to the kind of information he sends out.

I've seen the paid subscriber sections and it's nothing groundbreaking. I wouldn't/don't pay for it.

SBF used his altruism to cover up fraud. If the SA benchmarks were fraudulent, that would be a big deal. If he's just "in bed with the AI companies", like, that's a big part of the reason it's such a popular blog?

Also, I don't really have to think that Dylan is a good guy, and I certainly wasn't the only one that did not think SBF was a good guy. I get that it's really easy to call his implosion unsurprising after the fact, but it was truly unsurprising.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#386

Earlier quoted context omitted.

If we are applying Jevons paradox to this then the unit being consumed is not tokens but the inputs for token production - power, capex, something else. To draw an analogy to the steam engine, coal:electricity::mechanical-work:tokens. Jevons paradox does not talk about mechanical work becoming cheaper in the short term setting up a sort of rubber band of demand creating spiking prices for mechanical work. Compared to…

All Jevon’s paradox says is that as a resource becomes cheaper total consumption of that resource increases. It applies equally well to the inputs of token production as it does to the tokens themselves. The former would describe the effect the sellers into AI companies see (energy, GPU chips, RAM etc - if they lower their prices they’ll have more overall consumption) while the latter describes what the AI companies…

Jevons paradox states it might increase. It is not an ironclad law and there are many many cases where increasing efficiency wrt. a certain resource really will decrease the total consumption of that resource. Yes tokens are an input themselves, but this thread is discussing hardware that is more efficient at generating tokens. To increase the efficiency by which tokens are converted into some other product would require innovation in some other area - harnesses, the models themselves, better skill from the prompters, etc.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#387
post #385

Earlier quoted context omitted.

It isn't about bunk takes or not. It is about the motivation behind doing something. Their takes are fabricated in such a way as to drive clicks to their business, where they are printing money selling MNDA to the highest bidder. Dylan uses his influence as a service and it is borderline criminal. He just sued a whistleblower employee. It is so blatant, he even lives and works directly with people in power who feed h…

I dunno, it's a blog, so I'm not so worried about the motivation behind it besides how it biases their takes. I think being close to people that feed you information might be prerequisite to the kind of information he sends out. I've seen the paid subscriber sections and it's nothing groundbreaking. I wouldn't/don't pay for it. SBF used his altruism to cover up fraud. If the SA benchmarks were fraudulent, that would…

  > "I'm not so worried about the motivation behind it besides how it biases their takes."
Troi oi. Read what you wrote again.

Dylan is a grifting fraud. I wrote a too long document explaining a ton of examples and you're handwaving it away.

SA benchmarks for inferencemax? Yea... AMD / NVidia put their best engineers on tuning, just for the benchmarks. It isn't about serving up inference fast, it is about appearing better on the charts to sell more chips.

It is a popular blog because it is an influence service. That's the whole point. Write things that get clicks.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#388

The reliance on Deepseek and Kimi as the benchmarks from every chip maker from NVIDIA to OpenAI is a good tell of where things are heading. In the next couple of years, hopefully we will have systems at home for everyday use and corporations can buy bulk from providers.

As opposed to closed-source models? Benchmarks for GPT Sol wouldn’t be particularly meaningful, as no one else can run the benchmark, and we don’t know what the exact model specs are. Picking the best open source models is really the best they can do.

Fair point. I might be misreading based on my hopes :/

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#389

Earlier quoted context omitted.

If you were to remove the heat at a sufficient rate by, say, turning the lid into a heat exchanger, you would have a stable system.

That's the problem, removing heat at a sufficient rate. Of course it can be done, but the most efficient way (in terms of cost) is just open loop evaporation. I'm not a datacenter engineer, but I used to work in the ski industry. Snowmaking systems use vast quantities of compressed air. It works better if that air is cool. Blowing hot compressed air out of a snow cannon means the air temperature (wet bulb to be speci…

Oh like one of those scenic cone towers like on a nuclear power plant?

Iiuc youre saying: its more cost-efficient to waste water using evaporative cooling so thats what we'll get, not that a closed loop with a heat exchanger is technically infeasible?

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#390
post #85

I think they talked about this being general purpose chip but I would think that Anthropic/OpenAI are at the scale now they could bake LLM weights into chips themselves. For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough. While 2 years ago nothing was useful more than 1 year long, there are many older m…

"Baking in" a model into a chip is a bad idea because chips take 2 years to tape out and then you're stuck doing inference on llama 3 in 2026 when fable/sol are available. Every accelerator is a tradeoff between flexibility and performance and GPUs are already pareto-optimal

I think it would make sense to upload the weights like a firmware into the chip rather than baking the weights itself directly onto the chip. If there is an option of easily updating the firmware from time to time it would work well.
Post reply on HN