Live data from Hacker News

Furiosa: 3.5x efficiency over H100s

furiosa.ai

61–70 of 165 posts

Re: Furiosa: 3.5x efficiency over H100s

#63
post #34

Earlier quoted context omitted.

I thought they were saying it was more efficient, as in tokens per watt. I didn’t see a direct comparison on that metric but maybe I didn’t look well enough.

Probably. Companies sell on efficiency when they know they lose on performance.

Right, but datacenters also very much operate on electrical cost so it’s not without merit.

Re: Furiosa: 3.5x efficiency over H100s

#64
post #55
post #7

I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…

> (at which point Intel typically fired their chip design group, hired everyone from AMD or whoever, and came out with Core or whatever) Didn't the Core architecture come from the Intel Pentium M Israeli team? https://en.wikipedia.org/wiki/Intel_Core_(microarchitecture)...

Yeah, that bit was pure snark - point was Intel’s gotten caught resting on their laurels a couple times when their architectures get a little long in the tooth, and often it’s existential enough that the team that pulls them out of it isn’t the one that put them in it.

Re: Furiosa: 3.5x efficiency over H100s

#65

Earlier quoted context omitted.

> The reason this matters is that LLMs are incredibly nifty often useful tools that are not AGI and also seem to be hitting a scaling wall I don't know who needs to hear this, but the real break through in AI that we have had is not LLMs, but generative AI. LLM is but one specific case. Furthermore, we have hit absolutely no walls. Go download a model from Jan 2024, another from Jan 2025 and one from this year and co…

> exponential Is this the second most abused english word (after 'literally')? > a model from Jan 2024, another from Jan 2025 and one from this year You literally can't tell the difference is 'exponential', quadratic, or whatever from three data points. Plus it's not my experience at all. Since Deepseek I haven't found models that one can run on consumer hardware get much better.

I’ve heard “orders of magnitude” used more than once to mean 4-5 times

Re: Furiosa: 3.5x efficiency over H100s

#66

Earlier quoted context omitted.

> The reason this matters is that LLMs are incredibly nifty often useful tools that are not AGI and also seem to be hitting a scaling wall I don't know who needs to hear this, but the real break through in AI that we have had is not LLMs, but generative AI. LLM is but one specific case. Furthermore, we have hit absolutely no walls. Go download a model from Jan 2024, another from Jan 2025 and one from this year and co…

There is a lot of talking past each other when discussing LLM performance. The average person whose typical use case is asking ChatGPT how long they need to boil an egg for hasn't seen improvements for 18 months. Meanwhile if you're super into something like local models for example the tangible improvements are without exaggeration happening almost monthly.

> The average person whose typical use case is asking ChatGPT how long they need to boil an egg for hasn't seen improvements for 18 months

I don’t think that’s true. I think both my mother and my mother-in-law would start to complain pretty quickly if they got pushed back to 4o. Change may have felt gradual, but I think that’s more a function of growing confidence in what they can expect the machine to do.

I also think “ask how long to boil an egg” is missing a lot here. Both use ChatGPT in place of Google for all sorts of shit these days, including plenty of stuff they shouldn’t (like: “will the city be doing garbage collection tomorrow?”). Both are pretty sharp women but neither is remotely technical.

Re: Furiosa: 3.5x efficiency over H100s

#67
post #7

I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…

> The reason this matters is that LLMs are incredibly nifty often useful tools that are not AGI and also seem to be hitting a scaling wall I don't know who needs to hear this, but the real break through in AI that we have had is not LLMs, but generative AI. LLM is but one specific case. Furthermore, we have hit absolutely no walls. Go download a model from Jan 2024, another from Jan 2025 and one from this year and co…

> Go download a model from Jan 2024, another from Jan 2025 and one from this year and compare.

I did. The old one is smarter.

(The newer ones are more verbose, though. If that impresses you, then you probably think members of parliament are geniuses.)

Re: Furiosa: 3.5x efficiency over H100s

#68

Got excited, then I saw it was for inference. yawns Seems like it would obviously be in TSMCs interest to give preferential taping to nvidia competitors, they benefit from having a less consolidated customer base bidding up their prices.

My best guess after dipping my toe into semiconductor fabrication a decade ago is that there is a mysterious guru in a cave under a volcano who decides which customers get access to which nodes at which prices.

Re: Furiosa: 3.5x efficiency over H100s

#69

Earlier quoted context omitted.

> The reason this matters is that LLMs are incredibly nifty often useful tools that are not AGI and also seem to be hitting a scaling wall I don't know who needs to hear this, but the real break through in AI that we have had is not LLMs, but generative AI. LLM is but one specific case. Furthermore, we have hit absolutely no walls. Go download a model from Jan 2024, another from Jan 2025 and one from this year and co…

There is a lot of talking past each other when discussing LLM performance. The average person whose typical use case is asking ChatGPT how long they need to boil an egg for hasn't seen improvements for 18 months. Meanwhile if you're super into something like local models for example the tangible improvements are without exaggeration happening almost monthly.

Random trivia are answered much better in my case.

Re: Furiosa: 3.5x efficiency over H100s

#70

Earlier quoted context omitted.

Based on conversations I've had with some people managing GPU's at scale in the datacenters, inference is an after thought. There is a gold rush for training right now, and that's where these massive clusters are being used. LLM's are probably a small fraction of the overall GPU compute in use right now. I suspect in the next 5 years we'll have full Hollywood movies being completely generated (at least the specialfx)…

Hollywood studios are breathing their last gasps now. Anyone will be able to use AI to create blockbuster type movies, Hollywood's moat around that is rapidly draining.

Anybody had the ability to write the next great novel for a while, but few succeed.
Post reply on HN