Live data from Hacker News

Furiosa: 3.5x efficiency over H100s

furiosa.ai

21–30 of 165 posts

Re: Furiosa: 3.5x efficiency over H100s

#21
post #10
post #4

What can it actually run? The fact their benchmark plot refers to Llama 3.1 8b signals to me that it's hand implemented for that model and likely can't run newer / larger models. Why else would you benchmark such an outdated model? Show me a benchmark for gpt-oss-120b or something similar to that.

Looking at their blog, they in fact ran gpt-oss-120b: https://furiosa.ai/blog/serving-gpt-oss-120b-at-5-8-ms-tpot-... I think Llama 3 focus mostly reflects demand. It may be hard to believe, but many people aren't even aware gpt-oss exists.

Now I'm interested ...

It still kind of makes the point that you are stuck with a very limited range of models that they are hand implementing. But at least it's a model I would actually use. Give me that in a box I can put in a standard data center with normal power supply and I'm definitely interested.

But I want to know the cost :-)

Re: Furiosa: 3.5x efficiency over H100s

#22
post #4

What can it actually run? The fact their benchmark plot refers to Llama 3.1 8b signals to me that it's hand implemented for that model and likely can't run newer / larger models. Why else would you benchmark such an outdated model? Show me a benchmark for gpt-oss-120b or something similar to that.

The fact that so many people are focusing solely on massive LLM models is an oversight by people that narrowly focusing on a tiny (but very lucrative) subdomain of AI applications.

Re: Furiosa: 3.5x efficiency over H100s

#23
post #10
post #4

What can it actually run? The fact their benchmark plot refers to Llama 3.1 8b signals to me that it's hand implemented for that model and likely can't run newer / larger models. Why else would you benchmark such an outdated model? Show me a benchmark for gpt-oss-120b or something similar to that.

Looking at their blog, they in fact ran gpt-oss-120b: https://furiosa.ai/blog/serving-gpt-oss-120b-at-5-8-ms-tpot-... I think Llama 3 focus mostly reflects demand. It may be hard to believe, but many people aren't even aware gpt-oss exists.

> we demonstrated running gpt-oss-120b on two RNGD chips [snip] at 5.8 ms per output token

That's 86 token/second/chip

By comparison, a H100 will do 2390 token/second/GPU

Am I comparing the wrong things somehow?

[1] https://inferencemax.semianalysis.com/

Re: Furiosa: 3.5x efficiency over H100s

#24
post #7

I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…

> The reason this matters is that LLMs are incredibly nifty often useful tools that are not AGI and also seem to be hitting a scaling wall I don't know who needs to hear this, but the real break through in AI that we have had is not LLMs, but generative AI. LLM is but one specific case. Furthermore, we have hit absolutely no walls. Go download a model from Jan 2024, another from Jan 2025 and one from this year and co…

>go download a model

GP was talking about commercially hosted LLMs running in datacenters, not free Chinese models.

Local is definitely still improving. That’s another reason the megacenter model (NVDA’s big line up forever plan) is either a financial catastrophe about to happen, or the biggest bailout ever.

Re: Furiosa: 3.5x efficiency over H100s

#25
post #18
post #7

I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…

> I am of the opinion that Nvidia's hit the wall with their current architecture Based on what? Their measured performance on things people care about keep going up, and their software stack keeps getting better and unlocking more performance on existing hardware Inference tests: https://inferencemax.semianalysis.com/ Training tests: https://www.lightly.ai/blog/nvidia-b200-vs-h100 https://newsletter.semianalysis.com/…

> Is that based just on the HN "it is lots of money so it can't possibly make sense" wisdom?

I mean the amount of money invested across just a handful of AI companies is currently staggering and their respective revenues are no where near where they need to be. That’s a valid reason to be skeptical. How many times have we seen speculative investment of this magnitude? It’s shifting entire municipal and state economies in the US.

OpenAI alone is currently projected to burn over $100 billion by what? 2028 or 2029? Forgot what I read the other day. Tens of billions a year. That is a hell of a gamble by investors.

Re: Furiosa: 3.5x efficiency over H100s

#26
post #18
post #7

I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…

> I am of the opinion that Nvidia's hit the wall with their current architecture Based on what? Their measured performance on things people care about keep going up, and their software stack keeps getting better and unlocking more performance on existing hardware Inference tests: https://inferencemax.semianalysis.com/ Training tests: https://www.lightly.ai/blog/nvidia-b200-vs-h100 https://newsletter.semianalysis.com/…

> Their measured performance on things people care about keep going up, and their software stack keeps getting better and unlocking more performance on existing hardware

I'm more concerned about fully-loaded dollars per token - including datacenter and power costs - rather than "does the chip go faster." If Nvidia couldn't make the chip go faster, there wouldn't be any debate, the question right now is "what is the cost of those improvements." I don't have the answer to that number, but the numbers going around for the costs of new datacenters doesn't give me a lot of optimism.

> Is that based just on the HN "it is lots of money so it can't possibly make sense" wisdom?

OpenAI has $1.15T in spend commitments over the next 10 years: https://tomtunguz.com/openai-hardware-spending-2025-2035/

As far as revenue, the released numbers from nearly anyone in this space are questionable - they're not public companies, we don't actually get to see inside the box. Torture the numbers right and they'll tell you anything you want to hear. What we _do_ get to see is, eg, Anthropic raising billions of dollars every ~3 months or so over the course of 2025. Maybe they're just that ambitious, but that's the kind of thing that makes me nervous.

Re: Furiosa: 3.5x efficiency over H100s

#27
post #23
post #10

Earlier quoted context omitted.

Looking at their blog, they in fact ran gpt-oss-120b: https://furiosa.ai/blog/serving-gpt-oss-120b-at-5-8-ms-tpot-... I think Llama 3 focus mostly reflects demand. It may be hard to believe, but many people aren't even aware gpt-oss exists.

> we demonstrated running gpt-oss-120b on two RNGD chips [snip] at 5.8 ms per output token That's 86 token/second/chip By comparison, a H100 will do 2390 token/second/GPU Am I comparing the wrong things somehow? [1] https://inferencemax.semianalysis.com/

I thought they were saying it was more efficient, as in tokens per watt. I didn’t see a direct comparison on that metric but maybe I didn’t look well enough.

Re: Furiosa: 3.5x efficiency over H100s

#28
post #9
post #7

I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…

> but nothing about the industry's finances add up right now. The acquisitions do. Remember Groq?

That may not be a good example because everyone is saying Groq isn't worth $20B.

Re: Furiosa: 3.5x efficiency over H100s

#29
post #28
post #9

Earlier quoted context omitted.

> but nothing about the industry's finances add up right now. The acquisitions do. Remember Groq?

That may not be a good example because everyone is saying Groq isn't worth $20B.

They were valued at $6.9B just three months before Nvidia bought them for $20B, triple the valuation. That figure seems to have been pulled out of thin air.

Re: Furiosa: 3.5x efficiency over H100s

#30
post #17
post #7

I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…

> I am of the opinion that Nvidia's hit the wall with their current architecture Not likely since TSMC has a new process with big gains. > The story with Intel Was that their fab couldn’t keep up not designs.

If Intel's original 10nm process and Cannon Lake had launched within Intel's original timeframe of 2016/17, it would have been class leading.

Instead, they couldn't get 10nm to work and launched one low-power SKU in 2018 that had almost half the die disabled, and stuck to 14nm from 2014-2021.

Post reply on HN