Live data from Hacker News

OpenAI Jalapeño: Better than Nvidia Blackwell

newsletter.semianalysis.com

271–280 of 389 posts

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#271
post #85

I think they talked about this being general purpose chip but I would think that Anthropic/OpenAI are at the scale now they could bake LLM weights into chips themselves. For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough. While 2 years ago nothing was useful more than 1 year long, there are many older m…

Yeah this feels like Altman not knowing engineering well enough to realize where the focus should be.

Like he is optimizing to keep providing a vanilla token factory when weighted chips are coming and local models will supplement.

My head canon is savvy chip execs will be etching architecture his OpenAI pioneered into their flagship products while trying to minimize how much foothold he can get in hardware. Murica done offshored it. Not ours to control.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#272

The reliance on Deepseek and Kimi as the benchmarks from every chip maker from NVIDIA to OpenAI is a good tell of where things are heading. In the next couple of years, hopefully we will have systems at home for everyday use and corporations can buy bulk from providers.

As opposed to closed-source models? Benchmarks for GPT Sol wouldn’t be particularly meaningful, as no one else can run the benchmark, and we don’t know what the exact model specs are.

Picking the best open source models is really the best they can do.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#274
post #238
post #161

Earlier quoted context omitted.

The metal masked ROM is basically only 2 metal/contact layers. It's not a full new design and tapeout. You could roll a new set of parameters every ~2-3months. It's not an architectural change. See statements below. https://www.eetimes.com/taalas-specializes-to-extremes-for-e... https://www.turingpost.com/p/taalas https://cambrian-ai.com/taalas-launches-hardcore-chip-with-i... Part of the key is that by moving even f…

But that means your different chips all have different sets of weights and are different generations. If none of that is baked into the chip as now then all the chips are running the latest weights every time. Even if you could ignore the stuff built into the chip when the time came, at that point you just wasted money on silicon that’s useless in 2-3 months.

Why would it be useless in 3 months?

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#275
post #85

I think they talked about this being general purpose chip but I would think that Anthropic/OpenAI are at the scale now they could bake LLM weights into chips themselves. For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough. While 2 years ago nothing was useful more than 1 year long, there are many older m…

In case anyone's interested in these niche startups like taalas, here are a few more: 1. https://matx.com/ 2. https://www.d-matrix.ai/ 3. https://www.etched.com/ 4. https://www.positron.ai/ 5. https://hyperaccel.ai/ 6. https://axelera.ai/ 7. https://www.enchargeai.com/ 8. https://furiosa.ai/

https://tensordyne.ai

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#276
post #238
post #161

Earlier quoted context omitted.

The metal masked ROM is basically only 2 metal/contact layers. It's not a full new design and tapeout. You could roll a new set of parameters every ~2-3months. It's not an architectural change. See statements below. https://www.eetimes.com/taalas-specializes-to-extremes-for-e... https://www.turingpost.com/p/taalas https://cambrian-ai.com/taalas-launches-hardcore-chip-with-i... Part of the key is that by moving even f…

But that means your different chips all have different sets of weights and are different generations. If none of that is baked into the chip as now then all the chips are running the latest weights every time. Even if you could ignore the stuff built into the chip when the time came, at that point you just wasted money on silicon that’s useless in 2-3 months.

Imagine Anthropic gives you Opus of 6 months ago but at much higher speeds and much lower cost (that they might or might not pass on).

Would you use it?

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#277
post #238
post #161

Earlier quoted context omitted.

The metal masked ROM is basically only 2 metal/contact layers. It's not a full new design and tapeout. You could roll a new set of parameters every ~2-3months. It's not an architectural change. See statements below. https://www.eetimes.com/taalas-specializes-to-extremes-for-e... https://www.turingpost.com/p/taalas https://cambrian-ai.com/taalas-launches-hardcore-chip-with-i... Part of the key is that by moving even f…

But that means your different chips all have different sets of weights and are different generations. If none of that is baked into the chip as now then all the chips are running the latest weights every time. Even if you could ignore the stuff built into the chip when the time came, at that point you just wasted money on silicon that’s useless in 2-3 months.

Since the current NVL72 are still at the ~15%/yr failure rate it's not clear your new data center is going to have half it's compute in 3 years. If you're still running H100s they draw >10x kWh/Mtoken as new designs. All of these systems become dated, but not all of them require entirely new infrastructure.

If a ROM rack running a near frontier agent model at >10ktoken/sec costs What these don't do is TRAINING, they only do INFERENCE, but they could do it pretty well.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#278
Funny semi analysis has credibility here of all places. The founder is well-known in the hardware circle to be a black-market information trader.

It works like this:

1. Founder befriends undergrad interns/graduate student interns, buys them gifts, invite them to dinner/yacht/house/vc parties etc, or pays them to write articles 2. Founder extracts insider information out of these interns 3. Founder sells this information to companies paying "consulting" fees

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#279

Earlier quoted context omitted.

> I know this is what Taalas was doing (acquired by AMD), here was their demo, https://chatjimmy.ai/ which is based on Llama 3.1 8B. It feels like this should start to happen soon. Taalas needed a giant chip (6nm) for an 8B model. At best you could use a more advanced node to try to put a MoE model across several chips working together, but you can’t have GPT Sol size models on a single chip like that.

> Taalas needed a giant chip (6nm) for an 8B model. You're phrasing it like it was kind of an inherent technical limitation with this kind of burning weights into silicon. Which is also not new, it goes back to the 1980s with fixed function digital signal processors and little linear regressions or hardware classifiers for industrial control systems, all are the same basic principle. It's just usually not worth it to…

> You're phrasing it like it was kind of an inherent technical limitation with this kind of burning weights into silicon.

It is.

As I said, they could have shrunk it with a smaller process node, but that's not at all close to what would be required for a GPT Sol size model.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#280

Warning, this is a long comment! (I’m trying to stick to sourced facts here and not overstate what they mean) I went down a rabbit hole after watching Dylan Patel on Dwarkesh today: https://www.youtube.com/watch?v=aV26V1UvkJw I was initially just surprised by how bullish Dylan is on OpenAI/Anthropic and how bearish he is on China, despite Chinese labs getting closer to US SOTA while offering inference at dramatically…

See my other comment about this person. I'm surprised people take this site seriously.
Post reply on HN