Live data from Hacker News

OpenAI Jalapeño: Better than Nvidia Blackwell

newsletter.semianalysis.com

341–350 of 390 posts

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#341
post #186

Earlier quoted context omitted.

> Humans are still 22x more efficient, which is not that far considering the rate of progress in this area. Based on a human output rate of 3.3 tok/s, which seems questionable as a means of comparison

I do believe that this is the trade off. We are more efficient but slower in terms of thinking (at the same level of intelligence). Some animals go much further in terms of that trade off, see https://en.wikipedia.org/wiki/Portia_(spider) for example.

Our brain is more like a MoE model activating only a few neurons for specific activities making it more efficient unlike a dense model activating all the params.

Also brain produces quality tokens @ 3.3 tps instead of fast generating hallucinated tokens by certain models. Thus MTP can produce low quality tokens at 2x speed.

Patience pays.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#343
lamb-labs.com built some cool stuff.... I'm inclined to believe we're going to see static model chips everywhere, but I'm no expert in software.

I was early at efabless.com well now chipfoundry.io - they've done about 800 chip tape outs.

They've been doing open source silicon tape outs for a decade plus.

Founder recently built this: https://nativechips.ai --- not involved but I'm inclined to believe it's the future of where the market is going. I'm skeptical of many of the AI chip design startups and whether they've actually taped out chips and how many and at what scale.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#344
post #311
post #85

I think they talked about this being general purpose chip but I would think that Anthropic/OpenAI are at the scale now they could bake LLM weights into chips themselves. For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough. While 2 years ago nothing was useful more than 1 year long, there are many older m…

Computer Science history has taught us that this decision was almost always wrong. Every time a company has spent resources doing this, a competitor innovated on the software and made the custom hardware irrelevant.

> Every time a company has spent resources doing this, a competitor innovated on the software and made the custom hardware irrelevant.

This often happened, though not always.

An important counterexample are 3D graphics cards, which basically put the OpenGL/DirectX fixed-function pipeline into silicon. It took a long time and many iterations to make the pipeline more programmable until the 3D graphics cards turned into modern GPUs.

Even today, GPUs live on as separate hardware in a computer instead of having become integrated into, say, the CPU. Intel's attempt to do something like this with the Larrabee project [1] was discontinued.

---

[1] https://en.wikipedia.org/w/index.php?title=Larrabee_(microar...

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#345
post #123

Earlier quoted context omitted.

Which in turn reminds me of Soundblaster audio cards! I suspect inference chips are closer to the GPU story than the Soundblaster story though. I remember one soundblaster card I bought came with a Lara Croft demo, that exploited the incredible immersion of real time dynamic reverb. Genuinely I think game audio took a few steps back from that heady era, the innovation in audio likely didn't sell as many cards as grap…

EAX was very powerful in its heyday, but it has died because of a thousand cuts. First we had to have the audio processor. Good EAX was available on top of the line cards, and they were not always cheap. Lower end chips got less features. Then we had to have the speaker setup to have the greatest sound, or needed to get a real 5.1 headphones, which were bulky and never provided the same fidelity. Then Microsoft chang…

Creative Labs also turned itself into a brand I actively avoid for various reasons.

In early 2000s for example, friend with something like a Sound Blaster Live, but couldn't use it anymore as they lost the drivers, their website only had downloads for driver updates, requiring you to still have a driver CD, so no more CD meant no way to get the drivers.

They had not done very much meaningful stuff since the release of EAX and used patents to prevent anyone else from competing with them. Onboard sound cards were by and large indistinguishable from a quality perspective as Creative Labs ones, but cheaper which Creative Labs combated mostly with lawyers as opposed to upping their game.

Some guy after being frustrated with a long-standing bug in drivers for their Creative Labs sound card, dug into the binaries and made a fix for it. During which they also discovered that you could simply flip a switch in the driver to unlock features only meant to be available on more expensive hardware. Creative Labs of course went straight to lawyers to shut them down.

My brother bought the Creative Labs WoW headset which would have its mic get progressively softer until he would leave and re-join the call/voice chat room. They never released a driver update to fix this.

By the time that Microsoft announced no more "hardware acceleration" for sound cards, I had zero sympathy for Creative Labs, I was already convinced that they made pretty shoddy hardware/software and were mostly riding on their reputation from the 90s and some patents they managed to get.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#346

This semi-analysis article reads a lot more like an OpenAI press release than a real analysis. And to be honest some of the statements seem like just straight up lies - they initially claim they were invited to benchmark it, and then half way down switch to claiming that OpenAI provided all the numbers. This really kind of sucks, because I want to read actual detailed nuanced and credible analysis of what's happening…

As Jensen said, these guys are childish. They might have some technical chops somewhere in the organisation but their communication style of hubris + meme is really grating. I can't take a research organisation seriously when they so clearly want attention on X.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#347
post #16

Earlier quoted context omitted.

semianalysis is pretty good

Are they? https://jon4hotaisle.substack.com/p/influence-as-a-service-s...

I've been reading them since before all of the AI hype, and I've always thought they're pretty good. You a few spicy takes with the overview/opinions/benchmarks. Better than semiaccurate.

The article you link says not a lot of criticisms with very many words, and the AI prose gets much worse towards the end, seemingly when the author also gave up on reading it. I am disappointing in the plagiarism though, especially of Ryan Smith.

I am much more interested in what you think of the site though vs your own experiences running a GPU cloud. I've seen your comments on it for a long time, it's super interesting. So if you think their takes are mostly bunk I'd consider it way more than this hot aisle guy.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#348
post #238

Earlier quoted context omitted.

But that means your different chips all have different sets of weights and are different generations. If none of that is baked into the chip as now then all the chips are running the latest weights every time. Even if you could ignore the stuff built into the chip when the time came, at that point you just wasted money on silicon that’s useless in 2-3 months.

Doesn't matter if the chip is 100x-1000x more efficient and faster, and you can just make a new one for new weights. Imagine being able to run GPT Sol at 1k tokens/second a year from now, at a 100x lower cost per token than now. Would that be useful? Or Qwen 3.8 27B at 10k tokens/second. The super long thinking that makes qwen so effective would take a couple of seconds.

> Imagine being able to run GPT Sol at 1k tokens/second a year from now, at a 100x lower cost per token than now. Would that be useful?

Given the current rate of change, it would be hard to guess either way. By some measures the cost at fixed quality score goes down vastly faster than that:

  A similar trend is evident in the cost of models scoring above 50% on GPQA, a substantially more challenging benchmark than MMLU. There, inference costs declined from $15 per million tokens in May 2024 to $0.12 per million tokens by December 2024 (Phi 4).
- https://hai.stanford.edu/assets/files/hai_ai-index-report-20...

15/0.12 -> factor of 125 cost reduction in 7 months.

But that may well be an extreme case. To show how broad the range is, another quote from the same publication:

  Depending on the task, LLM inference prices have fallen anywhere from 9 to 900 times per year.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#349

Earlier quoted context omitted.

EAX was very powerful in its heyday, but it has died because of a thousand cuts. First we had to have the audio processor. Good EAX was available on top of the line cards, and they were not always cheap. Lower end chips got less features. Then we had to have the speaker setup to have the greatest sound, or needed to get a real 5.1 headphones, which were bulky and never provided the same fidelity. Then Microsoft chang…

Fucking Microsoft killed off a lot of cool shit with potential during the 1990s

As you mention the 90's you're probably referring to their embrace, extend, extinguish strategy. This situation with the sound card hardware mixing was quite different and a lot later.

Windows Vista moved away from kernel mode drivers where possible to help address the issue of BSODs which were not uncommon on Windows XP. It turns out that Microsoft was incorrectly getting the blame for BSODs when it was actually the fault of buggy video and sound card drivers.

While Vista was regarded as a "bad" Windows, it laid most of the foundations which allowed Windows 7 to be regarded as "really good" (by Windows standards).

From a hardware perspective, by the time Windows 7 arrived pretty much all drivers had been updated for Vista and had their kinks worked out (mostly, I would still have my NVidia drivers crash on occasion, but because they were user mode now, instead of a BSOD the screen would go black for a few seconds after which my desktop would come back with a Windows pop-up saying something to the effect of "the video drivers crashed and had to be restarted", the real perpetrator now being blamed!). Windows 7 also did optimizations to make things less resource intensive and it also helped that PCs had more RAM compared to when Vista came out.

On the UAC front Windows 7 was also much better, they calibrated the UAC prompts to come up less often, but what also happened is a lot of the 3rd party software which was needlessly requiring admin rights for no good reason (except that it was badly written), had finally been fixed by the time Windows 7 was released.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#350
post #311

Earlier quoted context omitted.

Computer Science history has taught us that this decision was almost always wrong. Every time a company has spent resources doing this, a competitor innovated on the software and made the custom hardware irrelevant.

> Every time a company has spent resources doing this, a competitor innovated on the software and made the custom hardware irrelevant. This often happened, though not always. An important counterexample are 3D graphics cards, which basically put the OpenGL/DirectX fixed-function pipeline into silicon. It took a long time and many iterations to make the pipeline more programmable until the 3D graphics cards turned int…

It didn't happen for the original fixed function graphics though, CPUs kept eating their lunch every year or so with fun rendering techniques. We still use some of these techniques for 2d graphics.

Only when they became more general with shaders, and then added support for GPGPU, did it truly take off.

I think this generality is the lesson here, not the fact that GPUs are not CPUs.

Post reply on HN