Live data from Hacker News

OpenAI Jalapeño: Better than Nvidia Blackwell

newsletter.semianalysis.com

321–330 of 390 posts

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#324

Earlier quoted context omitted.

these GPUs make computation faster, I understand as of now maybe all the computation is used to generate yet another junk LinkedIn post or unnecessary RFC, but at some point this craze should settle and we will be left with powerful computation machines, which can be used for computing more useful things

The GPUs being paid for w/ billions in investment will be obsolete and e-waste in a few short years same as a Cray-2 was just a decade after its release. It's fine if you're one of the people selling shovels to gold miners for a while, but sucks to be building houses in the boom town?

I am not selling shovels, I just like to see when there is real competition on the market and I hate regulatory capture what Anthropic is trying to do for models, because they are scared of Chinese open weight models.

In terms of GPUs whole world with 8B people have only couple of viable options: Nvidia, AMD, Intel - and largest part of their inventory is going to enterprises to run those LLMs, and its impacting every consumer / hobby projects, like cheap phones, DIY electronics projects and so on.

I want to have more alternatives on the market

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#325

This semi-analysis article reads a lot more like an OpenAI press release than a real analysis. And to be honest some of the statements seem like just straight up lies - they initially claim they were invited to benchmark it, and then half way down switch to claiming that OpenAI provided all the numbers. This really kind of sucks, because I want to read actual detailed nuanced and credible analysis of what's happening…

The "better than" in the title should have given it away. You don't compare two sophisticated products that have different ecosystems and summarize your findings with such simplistic wording.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#326
post #303

Earlier quoted context omitted.

Imagine Anthropic gives you Opus of 6 months ago but at much higher speeds and much lower cost (that they might or might not pass on). Would you use it?

Yes, this would basically obviate Sonnet and Haiku. If you consider them 1 and 2 generations behind, respectively (that's not really what they are), you can still get a ton out of those older chips. Not to mention people still use older Opus versions happily. (In part because they don't like the new Opus but still, the cost effectiveness is a huge boon.)

GPT 5.6 Luna is already rather fast at 300 tokens per second, performs better than Opus 4.6 from what I know, and is very cheap.

I don't know why anyone would use Haiku.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#327

Earlier quoted context omitted.

The link you cited is not evaporative cooling and a pot of boiling water with a sealed lid on it is a pressure vessel which eventually explodes.

If you were to remove the heat at a sufficient rate by, say, turning the lid into a heat exchanger, you would have a stable system.

That's the problem, removing heat at a sufficient rate. Of course it can be done, but the most efficient way (in terms of cost) is just open loop evaporation.

I'm not a datacenter engineer, but I used to work in the ski industry. Snowmaking systems use vast quantities of compressed air. It works better if that air is cool. Blowing hot compressed air out of a snow cannon means the air temperature (wet bulb to be specific) needs to be colder to make snow.

Anyways, most air compression stations use water to cool the air, and then evaporative coolers to cool the water. The water is reused, but a ton (not sure of the percentage) is lost into the air. It's more or less a tower with a big fan on top, and water percolates down from the top, being cooled by the air as it goes. The water is then collected and pumped through the system again (but of course has to be always topped up to counteract what was lost to evaporation).

Anyways, long story short is it's most cost effective to just spray water into the air to cool water, as long as water is free/cheap.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#330

The article is a bit naive: > However, as previously mentioned, Jalapeño’s results are obtained without speculative decoding and Vera Rubin’s results use speculative decoding. Speculative decoding leads to a ~3-5x reduction in cost per token. When speculative decoding is implemented on Jalapeño, this will enable Jalapeño to serve tokens even more cost effectively. How much speculative decoding improves throughput is…

In the slides on twitter you can see Jalapeno CAN do speculative decoding. In fact they explicitly mention how compute is disaggregated 3 ways now: prefill, predict, decode, and how a huge Jalapeno advantage is that it uses dark sillicon to switch between these without having to move the KV cache which remains local.
Post reply on HN