Live data from Hacker News

Furiosa: 3.5x efficiency over H100s

furiosa.ai

31–40 of 165 posts

Re: Furiosa: 3.5x efficiency over H100s

#32

Earlier quoted context omitted.

> The reason this matters is that LLMs are incredibly nifty often useful tools that are not AGI and also seem to be hitting a scaling wall I don't know who needs to hear this, but the real break through in AI that we have had is not LLMs, but generative AI. LLM is but one specific case. Furthermore, we have hit absolutely no walls. Go download a model from Jan 2024, another from Jan 2025 and one from this year and co…

>go download a model GP was talking about commercially hosted LLMs running in datacenters, not free Chinese models. Local is definitely still improving. That’s another reason the megacenter model (NVDA’s big line up forever plan) is either a financial catastrophe about to happen, or the biggest bailout ever.

GPT 5.2 is an incredible leap over 5.1 / 5

Re: Furiosa: 3.5x efficiency over H100s

#33
post #26
post #18

Earlier quoted context omitted.

> I am of the opinion that Nvidia's hit the wall with their current architecture Based on what? Their measured performance on things people care about keep going up, and their software stack keeps getting better and unlocking more performance on existing hardware Inference tests: https://inferencemax.semianalysis.com/ Training tests: https://www.lightly.ai/blog/nvidia-b200-vs-h100 https://newsletter.semianalysis.com/…

> Their measured performance on things people care about keep going up, and their software stack keeps getting better and unlocking more performance on existing hardware I'm more concerned about fully-loaded dollars per token - including datacenter and power costs - rather than "does the chip go faster." If Nvidia couldn't make the chip go faster, there wouldn't be any debate, the question right now is "what is the c…

> OpenAI has $1.15T in spend commitments over the next 10 years

Yes, but those aren't contracted commitments, and we know some of them are equity swaps. For example "Microsoft ($250B Azure commitment)" from the footnote is an unknown amount of actual cash.

And I think it's fair to point out the other information in your link "OpenAI projects a 48% gross profit margin in 2025, improving to 70% by 2029."

Re: Furiosa: 3.5x efficiency over H100s

#34
post #23

Earlier quoted context omitted.

> we demonstrated running gpt-oss-120b on two RNGD chips [snip] at 5.8 ms per output token That's 86 token/second/chip By comparison, a H100 will do 2390 token/second/GPU Am I comparing the wrong things somehow? [1] https://inferencemax.semianalysis.com/

I thought they were saying it was more efficient, as in tokens per watt. I didn’t see a direct comparison on that metric but maybe I didn’t look well enough.

Probably. Companies sell on efficiency when they know they lose on performance.

Re: Furiosa: 3.5x efficiency over H100s

#35
post #29
post #28

Earlier quoted context omitted.

That may not be a good example because everyone is saying Groq isn't worth $20B.

They were valued at $6.9B just three months before Nvidia bought them for $20B, triple the valuation. That figure seems to have been pulled out of thin air.

Speaking generally: It makes sense for a acquisition price to be at a premium to valuation, between the dynamics where you have to convince leadership its better to be bought than to keep growing, and the expected risk posed by them as competition.

Most M&As arent done by value investors.

Re: Furiosa: 3.5x efficiency over H100s

#36
post #3

really weird graph where they're comparing to 3x H100 PCI-E which is a config I don't think anyone is using. they're trying to compare at iso-power? I just want to see their box vs a box of 8 h100s b/c that's what people would buy instead, and they can divide tokens and watts if that's the pitch.

Whats a more realistic config?

Re: Furiosa: 3.5x efficiency over H100s

#37
Is this from 2024? It mentions "With global data center demand at 60 GW in 2024"

Also, there is no mention of the latest-gen NVDA chips: 5 RNGD servers generate tokens at 3.5x the rate of a single H100 SXM at 15 kW. This is reduced to 1.5x if you instead use 3 H100 PCIe servers as the benchmark.

Re: Furiosa: 3.5x efficiency over H100s

#38
post #23
post #10

Earlier quoted context omitted.

Looking at their blog, they in fact ran gpt-oss-120b: https://furiosa.ai/blog/serving-gpt-oss-120b-at-5-8-ms-tpot-... I think Llama 3 focus mostly reflects demand. It may be hard to believe, but many people aren't even aware gpt-oss exists.

> we demonstrated running gpt-oss-120b on two RNGD chips [snip] at 5.8 ms per output token That's 86 token/second/chip By comparison, a H100 will do 2390 token/second/GPU Am I comparing the wrong things somehow? [1] https://inferencemax.semianalysis.com/

I think you are comparing latency with throughput. You can't take the inverse of latency to get throughput because concurrency is unknown. But then, RNGD result is probably with concurrency=1.

Re: Furiosa: 3.5x efficiency over H100s

#39
post #7

I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…

Based on conversations I've had with some people managing GPU's at scale in the datacenters, inference is an after thought. There is a gold rush for training right now, and that's where these massive clusters are being used.

LLM's are probably a small fraction of the overall GPU compute in use right now. I suspect in the next 5 years we'll have full Hollywood movies being completely generated (at least the specialfx) entirely by AI.

Re: Furiosa: 3.5x efficiency over H100s

#40
post #7

I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…

Remember that without real competition, Nvidia has little incentive to release something 16x faster when they could release something 2x faster 4 times.
Post reply on HN