Furiosa: 3.5x efficiency over H100s
31–40 of 165 posts
Re: Furiosa: 3.5x efficiency over H100s
#32Earlier quoted context omitted.
> The reason this matters is that LLMs are incredibly nifty often useful tools that are not AGI and also seem to be hitting a scaling wall I don't know who needs to hear this, but the real break through in AI that we have had is not LLMs, but generative AI. LLM is but one specific case. Furthermore, we have hit absolutely no walls. Go download a model from Jan 2024, another from Jan 2025 and one from this year and co…
>go download a model GP was talking about commercially hosted LLMs running in datacenters, not free Chinese models. Local is definitely still improving. That’s another reason the megacenter model (NVDA’s big line up forever plan) is either a financial catastrophe about to happen, or the biggest bailout ever.
Re: Furiosa: 3.5x efficiency over H100s
#33Earlier quoted context omitted.
> I am of the opinion that Nvidia's hit the wall with their current architecture Based on what? Their measured performance on things people care about keep going up, and their software stack keeps getting better and unlocking more performance on existing hardware Inference tests: https://inferencemax.semianalysis.com/ Training tests: https://www.lightly.ai/blog/nvidia-b200-vs-h100 https://newsletter.semianalysis.com/…
> Their measured performance on things people care about keep going up, and their software stack keeps getting better and unlocking more performance on existing hardware I'm more concerned about fully-loaded dollars per token - including datacenter and power costs - rather than "does the chip go faster." If Nvidia couldn't make the chip go faster, there wouldn't be any debate, the question right now is "what is the c…
Yes, but those aren't contracted commitments, and we know some of them are equity swaps. For example "Microsoft ($250B Azure commitment)" from the footnote is an unknown amount of actual cash.
And I think it's fair to point out the other information in your link "OpenAI projects a 48% gross profit margin in 2025, improving to 70% by 2029."
Re: Furiosa: 3.5x efficiency over H100s
#34Earlier quoted context omitted.
> we demonstrated running gpt-oss-120b on two RNGD chips [snip] at 5.8 ms per output token That's 86 token/second/chip By comparison, a H100 will do 2390 token/second/GPU Am I comparing the wrong things somehow? [1] https://inferencemax.semianalysis.com/
I thought they were saying it was more efficient, as in tokens per watt. I didn’t see a direct comparison on that metric but maybe I didn’t look well enough.
Re: Furiosa: 3.5x efficiency over H100s
#35Earlier quoted context omitted.
That may not be a good example because everyone is saying Groq isn't worth $20B.
They were valued at $6.9B just three months before Nvidia bought them for $20B, triple the valuation. That figure seems to have been pulled out of thin air.
Most M&As arent done by value investors.
Re: Furiosa: 3.5x efficiency over H100s
#36really weird graph where they're comparing to 3x H100 PCI-E which is a config I don't think anyone is using. they're trying to compare at iso-power? I just want to see their box vs a box of 8 h100s b/c that's what people would buy instead, and they can divide tokens and watts if that's the pitch.
Re: Furiosa: 3.5x efficiency over H100s
#37Also, there is no mention of the latest-gen NVDA chips: 5 RNGD servers generate tokens at 3.5x the rate of a single H100 SXM at 15 kW. This is reduced to 1.5x if you instead use 3 H100 PCIe servers as the benchmark.
Re: Furiosa: 3.5x efficiency over H100s
#38Earlier quoted context omitted.
Looking at their blog, they in fact ran gpt-oss-120b: https://furiosa.ai/blog/serving-gpt-oss-120b-at-5-8-ms-tpot-... I think Llama 3 focus mostly reflects demand. It may be hard to believe, but many people aren't even aware gpt-oss exists.
> we demonstrated running gpt-oss-120b on two RNGD chips [snip] at 5.8 ms per output token That's 86 token/second/chip By comparison, a H100 will do 2390 token/second/GPU Am I comparing the wrong things somehow? [1] https://inferencemax.semianalysis.com/
Re: Furiosa: 3.5x efficiency over H100s
#39I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…
LLM's are probably a small fraction of the overall GPU compute in use right now. I suspect in the next 5 years we'll have full Hollywood movies being completely generated (at least the specialfx) entirely by AI.
Re: Furiosa: 3.5x efficiency over H100s
#40I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…