Live data from Hacker News

OpenAI Jalapeño: Better than Nvidia Blackwell

newsletter.semianalysis.com

181–190 of 389 posts

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#182
post #85

I think they talked about this being general purpose chip but I would think that Anthropic/OpenAI are at the scale now they could bake LLM weights into chips themselves. For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough. While 2 years ago nothing was useful more than 1 year long, there are many older m…

> I know this is what Taalas was doing (acquired by AMD), here was their demo, https://chatjimmy.ai/ which is based on Llama 3.1 8B. It feels like this should start to happen soon.

Taalas needed a giant chip (6nm) for an 8B model.

At best you could use a more advanced node to try to put a MoE model across several chips working together, but you can’t have GPT Sol size models on a single chip like that.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#184

I love how now you have to consider the possible s** posting motivation behind analysis of a trillion dollar industry being conducted at a world-class level by a bunch of ex Reddit and 4Chan adjacent mods -- it's one of the best stories in AI that SemiAnalysis is not cut from the same cloth as Gartner McKinsey et al

> McKinsey

lol. lmao even.

Have you seen the quality of their output? I'd take Claude or ChatGPT Free Tier over advice from McKinsey these days.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#185
The reliance on Deepseek and Kimi as the benchmarks from every chip maker from NVIDIA to OpenAI is a good tell of where things are heading. In the next couple of years, hopefully we will have systems at home for everyday use and corporations can buy bulk from providers.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#186

I hadn't seen the token/Joules comparison with human speech before. Humans are still 22x more efficient, which is not that far considering the rate of progress in this area.

> Humans are still 22x more efficient, which is not that far considering the rate of progress in this area.

Based on a human output rate of 3.3 tok/s, which seems questionable as a means of comparison

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#187

I hadn't seen the token/Joules comparison with human speech before. Humans are still 22x more efficient, which is not that far considering the rate of progress in this area.

I wonder how that stacks up if you consider all the time you have to keep the body alive when it’s not actively producing “tokens”.

Careful, let's not put the whole matrix into stasis outside of business hours.

Productivity is not the only reason to let these meatbags burn oxygen.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#188

Earlier quoted context omitted.

We should be mindful of the context that many of these providers VERY likely have been selling their subscriptions at a substantial loss So as much as i agree “more profits to stakeholders screw the customer”, i think its more of an emergency to get to profitability before the music stops.

> We should be mindful of the context that many of these providers VERY likely have been selling their subscriptions at a substantial loss. what makes you think this?

He's a subscription truther. There's loads of them. OpenAI's profit increases with each subscription that is cancelled. Pretty soon they'll have more profit than God.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#189

Earlier quoted context omitted.

I can totally see how ternary would work from a physical implementation perspective but I have a really hard time visualizing anything using base-e, can you explain how such a thing would work in practice?

No it's impossible. But it would be optimal!

Ah, the spherical cow of number bases :) Thanks for the response, that saved me a sleepless night.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#190
post #85

I think they talked about this being general purpose chip but I would think that Anthropic/OpenAI are at the scale now they could bake LLM weights into chips themselves. For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough. While 2 years ago nothing was useful more than 1 year long, there are many older m…

"Baking in" a model into a chip is a bad idea because chips take 2 years to tape out and then you're stuck doing inference on llama 3 in 2026 when fable/sol are available. Every accelerator is a tradeoff between flexibility and performance and GPUs are already pareto-optimal
Post reply on HN