Live data from Hacker News

Furiosa: 3.5x efficiency over H100s

furiosa.ai

71–80 of 165 posts

Re: Furiosa: 3.5x efficiency over H100s

#71
post #7

I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…

You’re right but Nvidia enjoys an important advantage Intel had always used to mask their sloppy design work: the supply chain. You simply can’t source HBMs at scale because Nvidia bought everything, TSMC N3 is likewise fully booked and between Apple and Nvidia their 18A is probably already far gone and if you want to connect your artisanal inference hardware together then congratulations, Nvidia is the leader here t…

For instance, I believe Callcenters are in big trouble, and so are specialized contractors (like those prepping for an SOC submission etc).

It is, however, actually funny how bad e.g. the amazon chatbot (Rufus) is on amazon.com. When asked where a particular CC charge comes from, it does all sorts of SQL queries into my account, but it can't be bothered to give me the link to the actual charges (the page exists and solves the problem trivially).

So, maybe, the callcenter troubles will take some time to materialize.

Re: Furiosa: 3.5x efficiency over H100s

#72
post #7

I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…

> and the only way to make the economics of, eg, a Blackwell-powered datacenter make sense is to assume that the entire economy is going to be running on it, as opposed to some useful tools and some improved interfaces.

And I'm still convinced we're not paying real prices anywhere. Everyone is still trying to get market share so the prices are going to go up when this all needs to sustain itself. At that point, which use cases become too expensive and does that shrink it's applicability ?

Re: Furiosa: 3.5x efficiency over H100s

#73

Earlier quoted context omitted.

Based on conversations I've had with some people managing GPU's at scale in the datacenters, inference is an after thought. There is a gold rush for training right now, and that's where these massive clusters are being used. LLM's are probably a small fraction of the overall GPU compute in use right now. I suspect in the next 5 years we'll have full Hollywood movies being completely generated (at least the specialfx)…

Hollywood studios are breathing their last gasps now. Anyone will be able to use AI to create blockbuster type movies, Hollywood's moat around that is rapidly draining.

Have you....used any of the video generators? Nothing they create make any goddamn sense, they're a step above those fake acid trip simulators.

Re: Furiosa: 3.5x efficiency over H100s

#74
post #7

I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…

> nothing about the industry's finances add up right now Nothing about the industry’s finances, or about Anthropic and OpenAI’s finances? I look at the list of providers on OpenRouter for open models, and I don’t believe all of them are losing money. FWIW Anthropic claims (iirc) that they don’t lose money on inference. So I don’t think the industry or the model of selling inference is what’s in trouble there. I am mu…

It's become rather clear from the local LLM communities catching up that there is no moat. Everyone is still just barely figuring out how this nifty data structures produce such a powerful emergent behavior, there isn't any truly secret sauce yet.

Re: Furiosa: 3.5x efficiency over H100s

#75
post #4

What can it actually run? The fact their benchmark plot refers to Llama 3.1 8b signals to me that it's hand implemented for that model and likely can't run newer / larger models. Why else would you benchmark such an outdated model? Show me a benchmark for gpt-oss-120b or something similar to that.

The fact that so many people are focusing solely on massive LLM models is an oversight by people that narrowly focusing on a tiny (but very lucrative) subdomain of AI applications.

Namely killing people or surveiling people, dealers choice.

Re: Furiosa: 3.5x efficiency over H100s

#76
post #18
post #7

I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…

> I am of the opinion that Nvidia's hit the wall with their current architecture Based on what? Their measured performance on things people care about keep going up, and their software stack keeps getting better and unlocking more performance on existing hardware Inference tests: https://inferencemax.semianalysis.com/ Training tests: https://www.lightly.ai/blog/nvidia-b200-vs-h100 https://newsletter.semianalysis.com/…

>Further, I'm sure most people heard the mention of an unnamed enterprise paying Anthropic $5000/month per developer on inference

Companies have wasted more money on dumber things so spending isn't a good measure.

And what about the countless other AI companies? Anthropic has one of the top models for coding so that's like saying there ins't a problem pre dot com bubble because Amazon is doing fine.

The real effects of AI is measured in rising profit of the customers of those AI companies otherwise you're looking at the shovel sellers

Re: Furiosa: 3.5x efficiency over H100s

#78
post #73

Earlier quoted context omitted.

Hollywood studios are breathing their last gasps now. Anyone will be able to use AI to create blockbuster type movies, Hollywood's moat around that is rapidly draining.

Have you....used any of the video generators? Nothing they create make any goddamn sense, they're a step above those fake acid trip simulators.

> Nothing they create make any goddamn sense,

I wouldn’t be that dismissive. Some have managed to make impressive things with them (although nothing close to an actual movie, even a short).

https://www.youtube.com/watch?v=ET7Y1nNMXmA

A bit older: https://www.youtube.com/watch?v=8OOpYvxKhtY

Compared to two years ago: https://www.youtube.com/watch?v=LHeCTfQOQcs

Re: Furiosa: 3.5x efficiency over H100s

#79
post #34

Earlier quoted context omitted.

I thought they were saying it was more efficient, as in tokens per watt. I didn’t see a direct comparison on that metric but maybe I didn’t look well enough.

Probably. Companies sell on efficiency when they know they lose on performance.

If you have an efficient chip you can just have more of them and come out ahead. This isn't a CPU where single core performance is all that important.

Re: Furiosa: 3.5x efficiency over H100s

#80
post #26
post #18

Earlier quoted context omitted.

> I am of the opinion that Nvidia's hit the wall with their current architecture Based on what? Their measured performance on things people care about keep going up, and their software stack keeps getting better and unlocking more performance on existing hardware Inference tests: https://inferencemax.semianalysis.com/ Training tests: https://www.lightly.ai/blog/nvidia-b200-vs-h100 https://newsletter.semianalysis.com/…

> Their measured performance on things people care about keep going up, and their software stack keeps getting better and unlocking more performance on existing hardware I'm more concerned about fully-loaded dollars per token - including datacenter and power costs - rather than "does the chip go faster." If Nvidia couldn't make the chip go faster, there wouldn't be any debate, the question right now is "what is the c…

GPUs are supply constrained and price isn't declining that fast so why do you expect the token price price to decrease. I think the supply issue will resolve in 1-2 years as now they have good prediction of how fast the market would grow.

Nvidia is literally selling GPUs with 90% profit margin and still everything is out of stock, which is unheard of before.

Post reply on HN