Earlier quoted context omitted.
I would assume asic based llm would work really well. Why did it not go well?
https://chatjimmy.ai/ runs Llama 3.1-8B on an ASIC as a demo by https://taalas.com/ I believe. That's quite a few parameters shy of today's trillion-weight behemoths, but it is fast.
OpenAI Jalapeño: Better than Nvidia Blackwell
251–260 of 389 posts
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#252I think they talked about this being general purpose chip but I would think that Anthropic/OpenAI are at the scale now they could bake LLM weights into chips themselves. For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough. While 2 years ago nothing was useful more than 1 year long, there are many older m…
Probably! But not viable yet; the chips would be about a year behind SOTA. Note the ~16 months that the article quotes as being insanely fast to get this chip to tape-out (read: start producing). We'll have to bootstrap our way there: AI is actively being used to get us closer to viable lead times for this. Unfortunately, there's some real physical constraints: IIRC, manufacturing a wafer takes on the order of a mont…
This way newly post-trained model can be loaded and served the same day.
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#253I love how now you have to consider the possible s** posting motivation behind analysis of a trillion dollar industry being conducted at a world-class level by a bunch of ex Reddit and 4Chan adjacent mods -- it's one of the best stories in AI that SemiAnalysis is not cut from the same cloth as Gartner McKinsey et al
semianalysis is pretty good
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#254Earlier quoted context omitted.
This may just be a classic case of Jevons paradox: https://en.wikipedia.org/wiki/Jevons_paradox In short, better hardware will drive down token cost in the near-term, but will drive up the demand for tokens as it gets cheap enough for other sectors to start to use it heavily. It comes from steam engines where economists originally thought that coal demand would plummet with more efficient engines, but it actually jus…
If we are applying Jevons paradox to this then the unit being consumed is not tokens but the inputs for token production - power, capex, something else. To draw an analogy to the steam engine, coal:electricity::mechanical-work:tokens. Jevons paradox does not talk about mechanical work becoming cheaper in the short term setting up a sort of rubber band of demand creating spiking prices for mechanical work. Compared to…
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#255Earlier quoted context omitted.
If you think tanking Trillions in investments, warming the earth and increasion ocean water levels, creating water shortages and brown-outs is "good for all of us" - well, the rest of us beg to differ.
these GPUs make computation faster, I understand as of now maybe all the computation is used to generate yet another junk LinkedIn post or unnecessary RFC, but at some point this craze should settle and we will be left with powerful computation machines, which can be used for computing more useful things
It's fine if you're one of the people selling shovels to gold miners for a while, but sucks to be building houses in the boom town?
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#256Earlier quoted context omitted.
Depends on compute capacity. If we become supply constrained on tokens, then prices will necessarily go up.
no they dont because inference stacks are getting more efficient and models are getting more intelligent per parameter.
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#257They also fawn over the chip’s TDP when all other chips have to support 16 bit floating point and thus must run much hotter.
They make the classic mistake of equating max TDP with in-use-watts, and praise this magnificent (fictitious) performance per watt at FP8 with other chips’ max-TDP at FP16, which draw twice the power.
Evidence that the IPO can’t be far away.
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#258These nascent inference chip efforts are reminding me of the early 3dfx / riva / mach / powervr days. Will be interesting to see if inference chips are here to stay and, if so, who the eventual dominant player(s) will be
Which in turn reminds me of Soundblaster audio cards! I suspect inference chips are closer to the GPU story than the Soundblaster story though. I remember one soundblaster card I bought came with a Lara Croft demo, that exploited the incredible immersion of real time dynamic reverb. Genuinely I think game audio took a few steps back from that heady era, the innovation in audio likely didn't sell as many cards as grap…
It wasn't really the cool reverb effects or wave tables, though those were a nice bonus. It was just "I can tell my computer to make sound and it actually makes sound without days of troubleshooting."
Granted, similar things could be said about 3dfx. It's was a 3D card with drivers that actually worked.
And then there's the obvious "sound blasters and voodoos go in my computer, jalapeno goes in someone else's computer" thing.
Re: OpenAI Jalapeño: Better than Nvidia Blackwell
#259I think they talked about this being general purpose chip but I would think that Anthropic/OpenAI are at the scale now they could bake LLM weights into chips themselves. For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough. While 2 years ago nothing was useful more than 1 year long, there are many older m…