Funny semi analysis has credibility here of all places. The founder is well-known in the hardware circle to be a black-market information trader. It works like this: 1. Founder befriends undergrad interns/graduate student interns, buys them gifts, invite them to dinner/yacht/house/vc parties etc, or pays them to write articles 2. Founder extracts insider information out of these interns 3. Founder sells this informat…
OpenAI Jalapeño: Better than Nvidia Blackwell
371–380 of 390 posts
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#372Earlier quoted context omitted.
Are they? https://jon4hotaisle.substack.com/p/influence-as-a-service-s...
I've been reading them since before all of the AI hype, and I've always thought they're pretty good. You a few spicy takes with the overview/opinions/benchmarks. Better than semiaccurate. The article you link says not a lot of criticisms with very many words, and the AI prose gets much worse towards the end, seemingly when the author also gave up on reading it. I am disappointing in the plagiarism though, especially…
Their takes are fabricated in such a way as to drive clicks to their business, where they are printing money selling MNDA to the highest bidder.
Dylan uses his influence as a service and it is borderline criminal. He just sued a whistleblower employee. It is so blatant, he even lives and works directly with people in power who feed him information.
Kind of like how SBF used his altruism to cover up the fraud he was doing. Everyone thought he was a good guy, until they realized he wasn't.
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#373This semi-analysis article reads a lot more like an OpenAI press release than a real analysis. And to be honest some of the statements seem like just straight up lies - they initially claim they were invited to benchmark it, and then half way down switch to claiming that OpenAI provided all the numbers. This really kind of sucks, because I want to read actual detailed nuanced and credible analysis of what's happening…
> some of the statements seem like just straight up lies Why the angst ? I suspect this announcement punctured a lot of people's bubbles, and many are in disbelief and denial and hence the emotional reaction seen here. That a company which never designed chips could suddenly leapfrog the best in the industry. What many forget is that openAI and anthropic are in a unique position to own the end-user experience, and th…
Notably the only one of the chips that also has a strong inference focus (but not only) is AMD's M1950X, which trounces OpenAI's chip (20 vs 3.4 FP8 PFLOPS, 40 vs 13.4 FP4 PFLOPS, 23 vs 15 TB/sec memory bandwidth), although it does use a lot more power (2500 vs 700W).
Google's TPU (now 8th generation) is glaringly absent from the performance comparison.
At the end of the day what really matters is cost not performance since you can always just run more chips. Google are full stack optimized from chip to data center, and might be expected to have an advantage.
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#374Earlier quoted context omitted.
But that means your different chips all have different sets of weights and are different generations. If none of that is baked into the chip as now then all the chips are running the latest weights every time. Even if you could ignore the stuff built into the chip when the time came, at that point you just wasted money on silicon that’s useless in 2-3 months.
Since the current NVL72 are still at the ~15%/yr failure rate it's not clear your new data center is going to have half it's compute in 3 years. If you're still running H100s they draw >10x kWh/Mtoken as new designs. All of these systems become dated, but not all of them require entirely new infrastructure. If a ROM rack running a near frontier agent model at >10ktoken/sec costs What these don't do is TRAINING, they…
This would be a stupidly bad failure rate, basically the worst business decision you could make, especially if you're somehow on the hook for eating those losses (which seems to be the implication?). Is there a linkable source on this?
The only thing I could find is SemiAnalysis claims that 15% of Blackwells end up RMAed[1]. That appears to be a total failure rate, though, and if you're RMAing them, you're getting replacements. So that appears to be a pretty different state of things.
[1] https://www.dwarkesh.com/p/dylan-patel#:~:text=GPUs%20are%20...
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#375Earlier quoted context omitted.
Google rolled out TPUs in 2015. AWS released Inferentia and Trainium chips in 2020. If companies working on ML-specific chips was evidence that large transformer models have fully saturated their potential, the field would have been done circa GPT-2.
Neither of those companies core business model was serving llms
Chips are another axis for improvements in training and inference. Orgs large enough to explore the space have been doing it for at least a decade now. This is just a silly line of reasoning based on the faulty assumption that somehow, looking for increases in efficiency in training/inference means teams have reached some theoretical limit in model capability.
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#376Earlier quoted context omitted.
> At low concurrency scenarios, Jalapeño demonstrates remarkable interactivity, hitting over 700 tokens per sec per user at concurrency 1 on the DeepSeek R1 model. This is about the same rate you get out of Sol Ultraspeed. Why do you think extra tool calls like that would be so unthinkable? It'd run circles around this, especially if the problem can be split up among a live-collaborating agent swarm, so that it's not…
They are not. If the robot speech is a tool call, then for a fair comparison we need to take the tool call scaffolding (and probably the reasoning too) into account. So rather than a sentence of 10 tokens worth of speech being the output, the raw token output would be maybe 10x or 100x that. Even more if we consider the management of other aspects of the robot embodiment (or we reduce the brain's 20W number to whatev…
The real question is how expensive it is to coordinate between these different modalities, and I really don't see why it'd be all that much.
I half expect Boston Dynamics to show something like this off in Q4 or whatever.
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#377This semi-analysis article reads a lot more like an OpenAI press release than a real analysis. And to be honest some of the statements seem like just straight up lies - they initially claim they were invited to benchmark it, and then half way down switch to claiming that OpenAI provided all the numbers. This really kind of sucks, because I want to read actual detailed nuanced and credible analysis of what's happening…
Semianalysis is an AI hype organisation, not a serious, unbiased semiconductor reviewer/journalist like chipsandcheese nor a documentarian of the semiconductor industry like Asianonmetry (as it relates so strongly to the modern economies of Asia). If you see something from semianalysis, you can simply ignore it.
It's helpful to realize that Dylan (Semianalysis), Dwarkesh, Aschebrenner (the Situational Awareness guy) and Sholto Douglas (Anthropic) all share a house in SF, so what you are getting from any of them is the SF AI scene view of the world, which is interesting to know, but probably not the best predictor of how things are going to pan out.
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#378Earlier quoted context omitted.
The semianalysis people have scripts which incorrectly count their numerators and denominators all the time. All their benchmarks are flawed. It is such a slipshod operation and they charge exorbitant amounts of money for it.
Say more about this please
the problem is they're so cryptopilled, surprises are what they want. they don't look at surprises and think, "that's wrong." they look at surprises and double down!
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#379Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#380Earlier quoted context omitted.
Keep in mind that what previous work has done on a single chip with weights baked in was on a 8b parameter model. Sol is likely something in the 5T parameter range, perhaps higher. Serving the whole thing at BF16 is on the order of $3m in hardware just to serve it at all, and closer to $1-1.5m of hardware if it was being served as NVFP4. And power draw starting at high tens to low hundreds of kilowatts. Let's say a m…
ill give you that the way we talk about this ppl seem to think wed do this tomorrow, but in the 70s a kb of ram took an entire rack and tons of power also. Its seems equally plausible that we could go into a cycle of iterative refinement of baked model hardware that would end up in "personal ai" just like we got to personal computing.
Something like the next iteration of Cerebras hardware paired with HBM for KV cache + HBF for weights could be incredibly strong here and much more likely to see away to make into a product with some lifetime compared to "let's bake a old model into a very, very large number of custom chips, design all the interconnects from scratch, and pray". Maybe in 10-15 years once this all matures.
Now if that works out that means in '29/'30 we could easily see a run on NAND that's even worse than the current DRAM price issues, on top of the current increases. Fun times if that happens.