Live data from Hacker News

Furiosa: 3.5x efficiency over H100s

furiosa.ai

91–100 of 165 posts

Re: Furiosa: 3.5x efficiency over H100s

#91
post #3

really weird graph where they're comparing to 3x H100 PCI-E which is a config I don't think anyone is using. they're trying to compare at iso-power? I just want to see their box vs a box of 8 h100s b/c that's what people would buy instead, and they can divide tokens and watts if that's the pitch.

Whats a more realistic config?

8xGPUs per box. this has been the data center standard for the last 8ish years.

furthermore usually NVLink connected within the box (SXM instead of PCIe cards, although the physical data link is still PCIe.)

this is important because the daughter board provides PCIe switches which usually connect NVMe drives, NICs and GPUs together such that within that subcomplex there isn't any PCIe oversubscription.

since last year for a lot of providers the standard is the GB200 I'd argue.

Re: Furiosa: 3.5x efficiency over H100s

#92

Earlier quoted context omitted.

> exponential Is this the second most abused english word (after 'literally')? > a model from Jan 2024, another from Jan 2025 and one from this year You literally can't tell the difference is 'exponential', quadratic, or whatever from three data points. Plus it's not my experience at all. Since Deepseek I haven't found models that one can run on consumer hardware get much better.

I’ve heard “orders of magnitude” used more than once to mean 4-5 times

In binary 2x is one order of magnitude

Re: Furiosa: 3.5x efficiency over H100s

#94

How is this possible? Doing AI with "dual AMD EPYC processors". I thought you needed to have GPUs or something like that to do the matrix multiplications needed to train LLMs? Is that conventional wisdom wrong?

it uses own chip under the hood, see accelerator mentioned in spec.

Re: Furiosa: 3.5x efficiency over H100s

#95
post #79
post #34

Earlier quoted context omitted.

Probably. Companies sell on efficiency when they know they lose on performance.

If you have an efficient chip you can just have more of them and come out ahead. This isn't a CPU where single core performance is all that important.

Eh if there's a human on the other side single stream performance is going to matter to them.

Re: Furiosa: 3.5x efficiency over H100s

#96

Is it reasonable for me not to be able to read a single word of a text-based blog post because I don't have WebGL enabled?

you are not the target audience

whatever runs on typical investor/C-suite laptops and phones (so new iPhone/MacBook with "stock" Safari, maybe in corporate some cursed Windows setup with Chrome) is okay, and obviously they need to maxx out the glitter, it's the 2020s

Re: Furiosa: 3.5x efficiency over H100s

#97
post #7

I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…

What about TPUs? They are more efficient than nvidia GPUs, a huge amount of inference is done with them, and while they are not literally being sold to the public, the whole technology should be influencing the next steps of Nvidia just like AMD influenced Intel

My understanding is all of Google's AI is trained and run on quite old but well designed TPUs. For a while the issue was that developing these AI models still needed flexibility and customised hardware like TPUs couldn't accomodate that.

Now that the model architecture has settled into something a bit more predictable, I wouldn't be surprised if we saw a little more specialisation in the hardware.

Re: Furiosa: 3.5x efficiency over H100s

#98
post #7

I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…

> which is going to blow out the economics on inference

At this point, I don't even think they do the envelope math anymore. However much money investors will be duped into giving them, that's what they'll spend on compute. Just gotta stay alive until the IPO!

Re: Furiosa: 3.5x efficiency over H100s

#99
post #96

Is it reasonable for me not to be able to read a single word of a text-based blog post because I don't have WebGL enabled?

you are not the target audience whatever runs on typical investor/C-suite laptops and phones (so new iPhone/MacBook with "stock" Safari, maybe in corporate some cursed Windows setup with Chrome) is okay, and obviously they need to maxx out the glitter, it's the 2020s

I know people with iphones 17 pro who do not have webgl enabled for sanitary infosec reasons:)

probably they don't want this site to be scraped by LLMs which would be kinda ironic

Re: Furiosa: 3.5x efficiency over H100s

#100
I think it's actually really cool to focus on efficiency over just raw performance! The page for the cards themselves goes into more detail and has a pretty nice graph: https://furiosa.ai/rngd

You can see them admit that RNGD will be slower than a setup with H100 SXM cards, but at the same time the tokens per second per watt is way better!

Actually, I wonder how different that is from Cerebras chips, since they're very much optimized for speed and one would think that'd also affect the efficiency a whole bunch: https://www.cerebras.ai/

Post reply on HN