Live data from Hacker News

OpenAI Jalapeño: Better than Nvidia Blackwell

newsletter.semianalysis.com

281–290 of 389 posts

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#281
So out of all the inference only chips which ones can I buy?

The only report is a smartnic fpga from Alibaba where we take an onnx design and write our own. https://essenceia.github.io/projects/alibaba_cloud_fpga/

On my M2 Pro Mac Mini the ANE only allows 2 gigabytes compared to the Metal GPU which can use the system ram.

Currently playing with https://www.asus.com/motherboards-components/ai-accelerator/... which is a 4bit, 8bit and 16 bit ai inference chip with 8 gigabytes of ram.

The UGen300 has the Hailo-10H chipset.

The ASUS Store price for the ugen300-usb-8g costs $365.00 Canadian dollars.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#282

Funny semi analysis has credibility here of all places. The founder is well-known in the hardware circle to be a black-market information trader. It works like this: 1. Founder befriends undergrad interns/graduate student interns, buys them gifts, invite them to dinner/yacht/house/vc parties etc, or pays them to write articles 2. Founder extracts insider information out of these interns 3. Founder sells this informat…

So you're saying the information is reliable?

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#283
post #85

I think they talked about this being general purpose chip but I would think that Anthropic/OpenAI are at the scale now they could bake LLM weights into chips themselves. For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough. While 2 years ago nothing was useful more than 1 year long, there are many older m…

Probably! But not viable yet; the chips would be about a year behind SOTA. Note the ~16 months that the article quotes as being insanely fast to get this chip to tape-out (read: start producing). We'll have to bootstrap our way there: AI is actively being used to get us closer to viable lead times for this. Unfortunately, there's some real physical constraints: IIRC, manufacturing a wafer takes on the order of a mont…

> Note the ~16 months that the article quotes as being insanely fast to get this chip to tape-out (read: start producing).

Yeah, with any luck it would put pressure on Nvidia to charge less, and not just to OpenAI. With a little more luck, we would see all the other players do the same thing, driving down the price of actual GPUs from GPU manufacturers.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#284
post #85

I think they talked about this being general purpose chip but I would think that Anthropic/OpenAI are at the scale now they could bake LLM weights into chips themselves. For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough. While 2 years ago nothing was useful more than 1 year long, there are many older m…

In case anyone's interested in these niche startups like taalas, here are a few more: 1. https://matx.com/ 2. https://www.d-matrix.ai/ 3. https://www.etched.com/ 4. https://www.positron.ai/ 5. https://hyperaccel.ai/ 6. https://axelera.ai/ 7. https://www.enchargeai.com/ 8. https://furiosa.ai/

https://tsavoritesi.com/

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#285
post #238

Earlier quoted context omitted.

But that means your different chips all have different sets of weights and are different generations. If none of that is baked into the chip as now then all the chips are running the latest weights every time. Even if you could ignore the stuff built into the chip when the time came, at that point you just wasted money on silicon that’s useless in 2-3 months.

Why would it be useless in 3 months?

Because on hacker news the only thing that matters is being in the current news cycle and not whether your business is profitable.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#286
post #238
post #161

Earlier quoted context omitted.

The metal masked ROM is basically only 2 metal/contact layers. It's not a full new design and tapeout. You could roll a new set of parameters every ~2-3months. It's not an architectural change. See statements below. https://www.eetimes.com/taalas-specializes-to-extremes-for-e... https://www.turingpost.com/p/taalas https://cambrian-ai.com/taalas-launches-hardcore-chip-with-i... Part of the key is that by moving even f…

But that means your different chips all have different sets of weights and are different generations. If none of that is baked into the chip as now then all the chips are running the latest weights every time. Even if you could ignore the stuff built into the chip when the time came, at that point you just wasted money on silicon that’s useless in 2-3 months.

Doesn't matter if the chip is 100x-1000x more efficient and faster, and you can just make a new one for new weights. Imagine being able to run GPT Sol at 1k tokens/second a year from now, at a 100x lower cost per token than now. Would that be useful? Or Qwen 3.8 27B at 10k tokens/second. The super long thinking that makes qwen so effective would take a couple of seconds.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#287
post #85

I think they talked about this being general purpose chip but I would think that Anthropic/OpenAI are at the scale now they could bake LLM weights into chips themselves. For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough. While 2 years ago nothing was useful more than 1 year long, there are many older m…

> For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough. but you trade updatability, which I don't think is worth it yet.

LoRAs are a parameter efficient way to update models. The fine tuning stopgap only has to work long enough to extend the lifespan of a model until the next chip comes out.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#288
post #209

Earlier quoted context omitted.

It depends when the good enough level hits. Pretty sure we are almost there for most common applications of AI.

That's only half the problem. OpenAI is contractually obligated, if you will, to believe that models will continue improving at an impressive rate for the foreseeable future (otherwise their valuation makes no sense). If you believe that, then you should expect to get Sol-level performance out of a Luna-cost model within six months or a year. If you have a system with the weights baked in, that means you're going to…

If a strange and quirky architecture can increase their margins, they could start becoming so wildly profitable that they don't care about their investor driven valuation anymore.

Re: OpenAI Jalapeño: Better than Nvidia Blackwell

#289

Earlier quoted context omitted.

Bots do that because other platforms remove or hide posts with bad words

Humans do it because they've been raised not to swear.

> because they've been raised to grant advertisers and their vile spawn more rights than humans
Post reply on HN