R@6.....111
OpenAI Jalapeño: Better than Nvidia Blackwell
261–270 of 390 posts
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#262I think they talked about this being general purpose chip but I would think that Anthropic/OpenAI are at the scale now they could bake LLM weights into chips themselves. For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough. While 2 years ago nothing was useful more than 1 year long, there are many older m…
> I know this is what Taalas was doing (acquired by AMD), here was their demo, https://chatjimmy.ai/ which is based on Llama 3.1 8B. It feels like this should start to happen soon. Taalas needed a giant chip (6nm) for an 8B model. At best you could use a more advanced node to try to put a MoE model across several chips working together, but you can’t have GPT Sol size models on a single chip like that.
You're phrasing it like it was kind of an inherent technical limitation with this kind of burning weights into silicon. Which is also not new, it goes back to the 1980s with fixed function digital signal processors and little linear regressions or hardware classifiers for industrial control systems, all are the same basic principle.
It's just usually not worth it to go super small process node, because most models people thought to turn into silicon were pretty small parameter sizes. We're talking 10-100 weight regression or at most 2-4k weight neural net, used in some instrument or factory equipment. You can do a decent MNIST OCR with a 4k weight neural net. For this, 180/130nm is fine.
Or you might think it's required with their special 4-bit as transistor thing (plausible). It's more that when you're experimenting and iterating, TSMC 6nm is their advertised path for rapid prototyping at cost for proof of concepts. And that's already in hot demand, while good luck if you're a startup trying to break in with 3/4nm as your first run.
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#263Earlier quoted context omitted.
> Humans are still 22x more efficient, which is not that far considering the rate of progress in this area. Based on a human output rate of 3.3 tok/s, which seems questionable as a means of comparison
I do believe that this is the trade off. We are more efficient but slower in terms of thinking (at the same level of intelligence). Some animals go much further in terms of that trade off, see https://en.wikipedia.org/wiki/Portia_(spider) for example.
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#264Earlier quoted context omitted.
"640k (token context) should be enough for anyone."
I know what you're saying, but modulo things like losing track of what year it is as time passes by, a current frontier model is going to continue to be useful for many tasks for many years, even moreso if it's 5-10x faster due to the chip architecture. It's not that it would be the best forever, it's that it would be useful for plenty long enough to be worthwhile, even if there was better stuff available. In exactly…
the youd have to buy a new one to get a better model is a FEATURE not a bug.
like if im apple... and i can put a sol level llm in an iphone, market it as privacy first you own your data personal assistant, integrate it all over the os... and then when there is a better model/siri make all the users buy a new phone... thats how they "win" ai.
the old standbys of better screens thinner cameras and batteries arent enough anymore. its basically tapped out. all modern phones are as thin as they need as big as they need as fast as they need and last all day on a battery...
apple needs a new number to up thing that people can actually feel/see. model generations could be it... every year faster, smarter, more capbilities and integrations.
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#265Earlier quoted context omitted.
Yes, they could also sell me GPT Sol 5.6 or 5.7 on a chip and I’d probably buy it. It’s a really really useful model for me, I’m not sure how much better for coding I need it to be. For most things I find Sol good enough with a small amount of coaxing around my tastes.
Keep in mind that what previous work has done on a single chip with weights baked in was on a 8b parameter model. Sol is likely something in the 5T parameter range, perhaps higher. Serving the whole thing at BF16 is on the order of $3m in hardware just to serve it at all, and closer to $1-1.5m of hardware if it was being served as NVFP4. And power draw starting at high tens to low hundreds of kilowatts. Let's say a m…
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#266I think they talked about this being general purpose chip but I would think that Anthropic/OpenAI are at the scale now they could bake LLM weights into chips themselves. For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough. While 2 years ago nothing was useful more than 1 year long, there are many older m…
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#267I think they talked about this being general purpose chip but I would think that Anthropic/OpenAI are at the scale now they could bake LLM weights into chips themselves. For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough. While 2 years ago nothing was useful more than 1 year long, there are many older m…
Probably! But not viable yet; the chips would be about a year behind SOTA. Note the ~16 months that the article quotes as being insanely fast to get this chip to tape-out (read: start producing). We'll have to bootstrap our way there: AI is actively being used to get us closer to viable lead times for this. Unfortunately, there's some real physical constraints: IIRC, manufacturing a wafer takes on the order of a mont…
Maybe! But it also doesn't require the rate of improvement to slow down. As long as some current model is eventually "good enough" for general use, it could still be a market-killer at a very low marginal price thanks to ASIC. Even if slower, much more expensive models are 10x better, that doesn't actually diminish the utility of the ASIC model, as long as it's "good enough".
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#268I think they talked about this being general purpose chip but I would think that Anthropic/OpenAI are at the scale now they could bake LLM weights into chips themselves. For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough. While 2 years ago nothing was useful more than 1 year long, there are many older m…
In case anyone's interested in these niche startups like taalas, here are a few more: 1. https://matx.com/ 2. https://www.d-matrix.ai/ 3. https://www.etched.com/ 4. https://www.positron.ai/ 5. https://hyperaccel.ai/ 6. https://axelera.ai/ 7. https://www.enchargeai.com/ 8. https://furiosa.ai/
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#269I hadn't seen the token/Joules comparison with human speech before. Humans are still 22x more efficient, which is not that far considering the rate of progress in this area.
But the true number is IMO far bigger: orders of magnitude greater if we think in terms of equivalent performance.
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#270Earlier quoted context omitted.
This may just be a classic case of Jevons paradox: https://en.wikipedia.org/wiki/Jevons_paradox In short, better hardware will drive down token cost in the near-term, but will drive up the demand for tokens as it gets cheap enough for other sectors to start to use it heavily. It comes from steam engines where economists originally thought that coal demand would plummet with more efficient engines, but it actually jus…
This is exactly what I see happening now. Codex keeps doing these usage resets. What do I do? Burn even more tokens than ever before. I know I'm not the only one.
(I switched to using local models as usage limits, api instability and the concept of paying per token stresses me out)