Live data from Hacker News

Furiosa: 3.5x efficiency over H100s

furiosa.ai

101–110 of 165 posts

Re: Furiosa: 3.5x efficiency over H100s

#101

Earlier quoted context omitted.

I’ve heard “orders of magnitude” used more than once to mean 4-5 times

In binary 2x is one order of magnitude

exactly!

I've been wondering about this for quite a while now. Why does everybody automatically assume that I'm using the decimal system when saying "orders of magnitude"?!

Re: Furiosa: 3.5x efficiency over H100s

#102

Got excited, then I saw it was for inference. yawns Seems like it would obviously be in TSMCs interest to give preferential taping to nvidia competitors, they benefit from having a less consolidated customer base bidding up their prices.

Everything is currently pointing towards inference being the main cost driver for LLMs in the future. Test-time-compute requires huge amounts of tokens in inference and makes providing frontier models as services unprofitable.

Anyone not under some kind of export restrictions can scrounge together some GPUs to train a frontier model (hell, even DeepSeek which is under these restrictions could) but providing a service that can compete with OpenAI et al. will prove to be quite costly. 3x improvements in inference are therefore nothing to sneeze at IMO.

Re: Furiosa: 3.5x efficiency over H100s

#103
post #7

I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…

What about TPUs? They are more efficient than nvidia GPUs, a huge amount of inference is done with them, and while they are not literally being sold to the public, the whole technology should be influencing the next steps of Nvidia just like AMD influenced Intel

TPUs can be more efficient, but are quite difficult to program for efficiently (difficult to saturate). That is why Google tends to sell TPU-services, rather than raw access to TPUs, so they can control the stack and get good utilization. GPUs are easier to work with.

I think the software side of the story is underestimated. Nvidia has a big moat there and huge community support.

Re: Furiosa: 3.5x efficiency over H100s

#104
For those wondering how this differs from Nvidia GPUs:

Nvidia = flexible, general-purpose GPUs that excel at training and mixed workloads. Furiosa = purpose-built inference ASICs that trade flexibility for much better cost, power efficiency, and predictable latency at scale.

Re: Furiosa: 3.5x efficiency over H100s

#105

Earlier quoted context omitted.

> But when you just look at it from an inference perspective, looking at these data centres like token factories makes sense. So if you ignore the majority of the costs, then it makes sense. Opus 4.5 was released on November 25, 2025. That is less than 2 months ago. When they stop training new models, then we can forget about training costs.

Inference costs scale linearly with usage. R&D expenses do not. That's not to mention that Dario Amodei has said that their models actually have a good return, even when accounting for training costs [0]. [0] https://youtu.be/GcqQ1ebBqkc?si=Vs2R4taIhj3uwIyj&t=1088

> Inference costs scale linearly with usage. R&D expenses do not.

Do we know this is true for AI?

Re: Furiosa: 3.5x efficiency over H100s

#106

Earlier quoted context omitted.

In binary 2x is one order of magnitude

exactly! I've been wondering about this for quite a while now. Why does everybody automatically assume that I'm using the decimal system when saying "orders of magnitude"?!

Because, as xkcd 169 says, communicating badly and then actung smug when you're misunderstood is not cleverness. "Orders of magnitude" refers to a decimal system in the vast majority of uses (I must admit I have no concrete data on this, but I can find plenty of references to it being base-10 and only a suggestion that it could be sometihng else).

Unless you've explicitly stated that you mean something else, people have no reason to think that you mean something else.

Re: Furiosa: 3.5x efficiency over H100s

#107

Earlier quoted context omitted.

Consensus seems to be that the labs are profitable on inference. They are only losing money on training and free users. The competition requiring them to spend that money on training and free users does complicate things. But when you just look at it from an inference perspective, looking at these data centres like token factories makes sense. I would definitely pay more to get faster inference of Opus 4.5, for examp…

> But when you just look at it from an inference perspective, looking at these data centres like token factories makes sense. So if you ignore the majority of the costs, then it makes sense. Opus 4.5 was released on November 25, 2025. That is less than 2 months ago. When they stop training new models, then we can forget about training costs.

I'm not taking a side here - I don't know enough - but it's an interesting line of reasoning.

So I'll ask, how is that any different than fabs? From what I understand R&D is absurd and upgrading to a new node is even more absurd. The resulting chips sell for chump change on a per unit basis (analogous to tokens). But somehow it all works out.

Well, sort of. The bleeding edge companies kept dropping out until you could count them on one hand at this point.

At first glance it seems like the analogy might fit?

Re: Furiosa: 3.5x efficiency over H100s

#108
post #79
post #34

Earlier quoted context omitted.

Probably. Companies sell on efficiency when they know they lose on performance.

If you have an efficient chip you can just have more of them and come out ahead. This isn't a CPU where single core performance is all that important.

Only if the price is right...

Re: Furiosa: 3.5x efficiency over H100s

#109

I think it's actually really cool to focus on efficiency over just raw performance! The page for the cards themselves goes into more detail and has a pretty nice graph: https://furiosa.ai/rngd You can see them admit that RNGD will be slower than a setup with H100 SXM cards, but at the same time the tokens per second per watt is way better! Actually, I wonder how different that is from Cerebras chips, since they're ve…

Having only 48GB of RAM per card seems low. The full server system with 8 cards barely has enough RAM to run modern large open models. And batching together user requests eats quite a lot of memory, too. Curious to see how these machines and cards are received by the market.

Re: Furiosa: 3.5x efficiency over H100s

#110
post #3

really weird graph where they're comparing to 3x H100 PCI-E which is a config I don't think anyone is using. they're trying to compare at iso-power? I just want to see their box vs a box of 8 h100s b/c that's what people would buy instead, and they can divide tokens and watts if that's the pitch.

It would also depend on the purchase cost and cooling infrastructure cost. If this costs what a 3x H100 box costs then it’s a fair comparison even if not a direct comparison to what customers currently buy.
Post reply on HN