Live data from Hacker News

OpenAI Jalapeño: Better than Nvidia Blackwell

newsletter.semianalysis.com

371–380 of 390 posts

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#371

Funny semi analysis has credibility here of all places. The founder is well-known in the hardware circle to be a black-market information trader. It works like this: 1. Founder befriends undergrad interns/graduate student interns, buys them gifts, invite them to dinner/yacht/house/vc parties etc, or pays them to write articles 2. Founder extracts insider information out of these interns 3. Founder sells this informat…

So this is like the government attacking journalists for surfacing damning information, when in reality it's the government doing bad shit? Don't go after the interns doing the leaking, it's just easier to malign the messenger.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#372
post #347

Earlier quoted context omitted.

Are they? https://jon4hotaisle.substack.com/p/influence-as-a-service-s...

I've been reading them since before all of the AI hype, and I've always thought they're pretty good. You a few spicy takes with the overview/opinions/benchmarks. Better than semiaccurate. The article you link says not a lot of criticisms with very many words, and the AI prose gets much worse towards the end, seemingly when the author also gave up on reading it. I am disappointing in the plagiarism though, especially…

It isn't about bunk takes or not. It is about the motivation behind doing something.

Their takes are fabricated in such a way as to drive clicks to their business, where they are printing money selling MNDA to the highest bidder.

Dylan uses his influence as a service and it is borderline criminal. He just sued a whistleblower employee. It is so blatant, he even lives and works directly with people in power who feed him information.

Kind of like how SBF used his altruism to cover up the fraud he was doing. Everyone thought he was a good guy, until they realized he wasn't.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#373

This semi-analysis article reads a lot more like an OpenAI press release than a real analysis. And to be honest some of the statements seem like just straight up lies - they initially claim they were invited to benchmark it, and then half way down switch to claiming that OpenAI provided all the numbers. This really kind of sucks, because I want to read actual detailed nuanced and credible analysis of what's happening…

> some of the statements seem like just straight up lies Why the angst ? I suspect this announcement punctured a lot of people's bubbles, and many are in disbelief and denial and hence the emotional reaction seen here. That a company which never designed chips could suddenly leapfrog the best in the industry. What many forget is that openAI and anthropic are in a unique position to own the end-user experience, and th…

A lot of the comparison is apples and oranges though - it seems that OpenAI's chip is targeting inference (FP8, FP4), while most of the chips it is being compared to are general purpose.

Notably the only one of the chips that also has a strong inference focus (but not only) is AMD's M1950X, which trounces OpenAI's chip (20 vs 3.4 FP8 PFLOPS, 40 vs 13.4 FP4 PFLOPS, 23 vs 15 TB/sec memory bandwidth), although it does use a lot more power (2500 vs 700W).

Google's TPU (now 8th generation) is glaringly absent from the performance comparison.

At the end of the day what really matters is cost not performance since you can always just run more chips. Google are full stack optimized from chip to data center, and might be expected to have an advantage.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#374
post #277
post #238

Earlier quoted context omitted.

But that means your different chips all have different sets of weights and are different generations. If none of that is baked into the chip as now then all the chips are running the latest weights every time. Even if you could ignore the stuff built into the chip when the time came, at that point you just wasted money on silicon that’s useless in 2-3 months.

Since the current NVL72 are still at the ~15%/yr failure rate it's not clear your new data center is going to have half it's compute in 3 years. If you're still running H100s they draw >10x kWh/Mtoken as new designs. All of these systems become dated, but not all of them require entirely new infrastructure. If a ROM rack running a near frontier agent model at >10ktoken/sec costs What these don't do is TRAINING, they…

> Since the current NVL72 are still at the ~15%/yr failure rate it's not clear your new data center is going to have half it's compute in 3 years

This would be a stupidly bad failure rate, basically the worst business decision you could make, especially if you're somehow on the hook for eating those losses (which seems to be the implication?). Is there a linkable source on this?

The only thing I could find is SemiAnalysis claims that 15% of Blackwells end up RMAed[1]. That appears to be a total failure rate, though, and if you're RMAing them, you're getting replacements. So that appears to be a pretty different state of things.

[1] https://www.dwarkesh.com/p/dylan-patel#:~:text=GPUs%20are%20...

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#375

Earlier quoted context omitted.

Google rolled out TPUs in 2015. AWS released Inferentia and Trainium chips in 2020. If companies working on ML-specific chips was evidence that large transformer models have fully saturated their potential, the field would have been done circa GPT-2.

Neither of those companies core business model was serving llms

What? Both of those companies absolutely serve LLMs, and both of them would love for serving LLMs to be an even bigger part of their business. Not only that, AWS is Anthropic's primary compute partner for training and inference. They literally use the newest generation of the Trainium chips I mentioned before: https://www.anthropic.com/news/anthropic-amazon-compute

Chips are another axis for improvements in training and inference. Orgs large enough to explore the space have been doing it for at least a decade now. This is just a silly line of reasoning based on the faulty assumption that somehow, looking for increases in efficiency in training/inference means teams have reached some theoretical limit in model capability.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#376

Earlier quoted context omitted.

> At low concurrency scenarios, Jalapeño demonstrates remarkable interactivity, hitting over 700 tokens per sec per user at concurrency 1 on the DeepSeek R1 model. This is about the same rate you get out of Sol Ultraspeed. Why do you think extra tool calls like that would be so unthinkable? It'd run circles around this, especially if the problem can be split up among a live-collaborating agent swarm, so that it's not…

They are not. If the robot speech is a tool call, then for a fair comparison we need to take the tool call scaffolding (and probably the reasoning too) into account. So rather than a sentence of 10 tokens worth of speech being the output, the raw token output would be maybe 10x or 100x that. Even more if we consider the management of other aspects of the robot embodiment (or we reduce the brain's 20W number to whatev…

But there are already voice models that do a reasonable job at a fraction of the throughput available?

The real question is how expensive it is to coordinate between these different modalities, and I really don't see why it'd be all that much.

I half expect Boston Dynamics to show something like this off in Q4 or whatever.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#377

This semi-analysis article reads a lot more like an OpenAI press release than a real analysis. And to be honest some of the statements seem like just straight up lies - they initially claim they were invited to benchmark it, and then half way down switch to claiming that OpenAI provided all the numbers. This really kind of sucks, because I want to read actual detailed nuanced and credible analysis of what's happening…

Semianalysis is an AI hype organisation, not a serious, unbiased semiconductor reviewer/journalist like chipsandcheese nor a documentarian of the semiconductor industry like Asianonmetry (as it relates so strongly to the modern economies of Asia). If you see something from semianalysis, you can simply ignore it.

Semianalysis seem to have useful information but increasingly crazy extrapolations of trends and future predictions.

It's helpful to realize that Dylan (Semianalysis), Dwarkesh, Aschebrenner (the Situational Awareness guy) and Sholto Douglas (Anthropic) all share a house in SF, so what you are getting from any of them is the SF AI scene view of the world, which is interesting to know, but probably not the best predictor of how things are going to pan out.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#378

Earlier quoted context omitted.

The semianalysis people have scripts which incorrectly count their numerators and denominators all the time. All their benchmarks are flawed. It is such a slipshod operation and they charge exorbitant amounts of money for it.

Say more about this please

for example, their people think that GB300s are twice as fast as B300s, when really their benchmarks just incorrectly divide GB300 instances on azure by 8 instead of 4, since they don't read or verify any of the code that executes their benchmarks.

the problem is they're so cryptopilled, surprises are what they want. they don't look at surprises and think, "that's wrong." they look at surprises and double down!

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#380

Earlier quoted context omitted.

Keep in mind that what previous work has done on a single chip with weights baked in was on a 8b parameter model. Sol is likely something in the 5T parameter range, perhaps higher. Serving the whole thing at BF16 is on the order of $3m in hardware just to serve it at all, and closer to $1-1.5m of hardware if it was being served as NVFP4. And power draw starting at high tens to low hundreds of kilowatts. Let's say a m…

ill give you that the way we talk about this ppl seem to think wed do this tomorrow, but in the 70s a kb of ram took an entire rack and tons of power also. Its seems equally plausible that we could go into a cycle of iterative refinement of baked model hardware that would end up in "personal ai" just like we got to personal computing.

Baked model hardware is not the next step in the chain here, in the next 2-3 years we might hopefully see some HBF (high bandwidth flash) hardware to try and get at same memory bandwidth today at a somewhat lower price point and much lower power dissipation.

Something like the next iteration of Cerebras hardware paired with HBM for KV cache + HBF for weights could be incredibly strong here and much more likely to see away to make into a product with some lifetime compared to "let's bake a old model into a very, very large number of custom chips, design all the interconnects from scratch, and pray". Maybe in 10-15 years once this all matures.

Now if that works out that means in '29/'30 we could easily see a run on NAND that's even worse than the current DRAM price issues, on top of the current increases. Fun times if that happens.

Post reply on HN