Live data from Hacker News

Furiosa: 3.5x efficiency over H100s

furiosa.ai

81–90 of 165 posts

Re: Furiosa: 3.5x efficiency over H100s

#81

Earlier quoted context omitted.

I am not someone who would ever be ever be considered an expert on factories/manufacturing of any kind, but my (insanely basic) understanding is that typically a “factory” making whatever widgets or doodads is outputting at a profit or has a clear path to profitability in order to pay off a loan/investment. They have debt, but they’re moving towards the black in a concrete, relatively predictable way - no one specula…

Consensus seems to be that the labs are profitable on inference. They are only losing money on training and free users. The competition requiring them to spend that money on training and free users does complicate things. But when you just look at it from an inference perspective, looking at these data centres like token factories makes sense. I would definitely pay more to get faster inference of Opus 4.5, for examp…

> But when you just look at it from an inference perspective, looking at these data centres like token factories makes sense.

So if you ignore the majority of the costs, then it makes sense.

Opus 4.5 was released on November 25, 2025. That is less than 2 months ago. When they stop training new models, then we can forget about training costs.

Re: Furiosa: 3.5x efficiency over H100s

#82
post #74

Earlier quoted context omitted.

> nothing about the industry's finances add up right now Nothing about the industry’s finances, or about Anthropic and OpenAI’s finances? I look at the list of providers on OpenRouter for open models, and I don’t believe all of them are losing money. FWIW Anthropic claims (iirc) that they don’t lose money on inference. So I don’t think the industry or the model of selling inference is what’s in trouble there. I am mu…

It's become rather clear from the local LLM communities catching up that there is no moat. Everyone is still just barely figuring out how this nifty data structures produce such a powerful emergent behavior, there isn't any truly secret sauce yet.

I’d argue there’s a _bit_ of secret sauce here, but the question is if there’s enough to justify valuations of the prop-AI firms, and that seems unlikely.

Re: Furiosa: 3.5x efficiency over H100s

#83
post #7

I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…

What about TPUs? They are more efficient than nvidia GPUs, a huge amount of inference is done with them, and while they are not literally being sold to the public, the whole technology should be influencing the next steps of Nvidia just like AMD influenced Intel

Re: Furiosa: 3.5x efficiency over H100s

#84
post #33
post #26

Earlier quoted context omitted.

> Their measured performance on things people care about keep going up, and their software stack keeps getting better and unlocking more performance on existing hardware I'm more concerned about fully-loaded dollars per token - including datacenter and power costs - rather than "does the chip go faster." If Nvidia couldn't make the chip go faster, there wouldn't be any debate, the question right now is "what is the c…

> OpenAI has $1.15T in spend commitments over the next 10 years Yes, but those aren't contracted commitments, and we know some of them are equity swaps. For example "Microsoft ($250B Azure commitment)" from the footnote is an unknown amount of actual cash. And I think it's fair to point out the other information in your link "OpenAI projects a 48% gross profit margin in 2025, improving to 70% by 2029."

> Yes, but those aren't contracted commitments, and we know some of them are equity swaps.

It's worse than not contracted. Nvidia said in their earnings call that their OpenAI commitment was "maybe".

Re: Furiosa: 3.5x efficiency over H100s

#85
post #18

Earlier quoted context omitted.

> I am of the opinion that Nvidia's hit the wall with their current architecture Based on what? Their measured performance on things people care about keep going up, and their software stack keeps getting better and unlocking more performance on existing hardware Inference tests: https://inferencemax.semianalysis.com/ Training tests: https://www.lightly.ai/blog/nvidia-b200-vs-h100 https://newsletter.semianalysis.com/…

> Is that based just on the HN "it is lots of money so it can't possibly make sense" wisdom? I mean the amount of money invested across just a handful of AI companies is currently staggering and their respective revenues are no where near where they need to be. That’s a valid reason to be skeptical. How many times have we seen speculative investment of this magnitude? It’s shifting entire municipal and state economie…

> I mean the amount of money invested across just a handful of AI companies is currently staggering and their respective revenues are no where near where they need to be. That’s a valid reason to be skeptical.

Yes and no. Some of it just claims to be "AI". Like the hyperscalers are building datacenters and ramping up but not all of it is "AI". The crypto bros have rebadged their data centers into "AI".

Re: Furiosa: 3.5x efficiency over H100s

#86
post #55
post #7

I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…

> (at which point Intel typically fired their chip design group, hired everyone from AMD or whoever, and came out with Core or whatever) Didn't the Core architecture come from the Intel Pentium M Israeli team? https://en.wikipedia.org/wiki/Intel_Core_(microarchitecture)...

Yes, and the newest Panther Lake too!

https://techtime.news/2025/10/10/intel-25/

Re: Furiosa: 3.5x efficiency over H100s

#87

Earlier quoted context omitted.

> The reason this matters is that LLMs are incredibly nifty often useful tools that are not AGI and also seem to be hitting a scaling wall I don't know who needs to hear this, but the real break through in AI that we have had is not LLMs, but generative AI. LLM is but one specific case. Furthermore, we have hit absolutely no walls. Go download a model from Jan 2024, another from Jan 2025 and one from this year and co…

> Go download a model from Jan 2024, another from Jan 2025 and one from this year and compare. I did. The old one is smarter. (The newer ones are more verbose, though. If that impresses you, then you probably think members of parliament are geniuses.)

Yeah agreed, there were some minor gains, but new releases are mostly benchmark overfit sycopanthic bullshit that are only better on paper and horrible to use. The more synthetic data they add the less world knowledge the model has and the more useless it becomes. But at least they can almost mimic a basic calculator now /s

For api models, OpenAI's releases have regularly not been an improvement for a long while now. Is sonnet 4.5 better than 3.5 outside pretentius agentic workflows it's been trained for? Basically impossible to tell, they make the same braindead mistakes sometimes.

Re: Furiosa: 3.5x efficiency over H100s

#88
post #7

I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…

What do I care if there's no profit in LLM's..

I just want to buy ddr5 and not pay an arm and a leg for my power bill!

Re: Furiosa: 3.5x efficiency over H100s

#89
post #33
post #26

Earlier quoted context omitted.

> Their measured performance on things people care about keep going up, and their software stack keeps getting better and unlocking more performance on existing hardware I'm more concerned about fully-loaded dollars per token - including datacenter and power costs - rather than "does the chip go faster." If Nvidia couldn't make the chip go faster, there wouldn't be any debate, the question right now is "what is the c…

> OpenAI has $1.15T in spend commitments over the next 10 years Yes, but those aren't contracted commitments, and we know some of them are equity swaps. For example "Microsoft ($250B Azure commitment)" from the footnote is an unknown amount of actual cash. And I think it's fair to point out the other information in your link "OpenAI projects a 48% gross profit margin in 2025, improving to 70% by 2029."

Sounds like the railway boom.. I mean bond scam's

Re: Furiosa: 3.5x efficiency over H100s

#90

Earlier quoted context omitted.

Consensus seems to be that the labs are profitable on inference. They are only losing money on training and free users. The competition requiring them to spend that money on training and free users does complicate things. But when you just look at it from an inference perspective, looking at these data centres like token factories makes sense. I would definitely pay more to get faster inference of Opus 4.5, for examp…

> But when you just look at it from an inference perspective, looking at these data centres like token factories makes sense. So if you ignore the majority of the costs, then it makes sense. Opus 4.5 was released on November 25, 2025. That is less than 2 months ago. When they stop training new models, then we can forget about training costs.

Inference costs scale linearly with usage. R&D expenses do not.

That's not to mention that Dario Amodei has said that their models actually have a good return, even when accounting for training costs [0].

[0] https://youtu.be/GcqQ1ebBqkc?si=Vs2R4taIhj3uwIyj&t=1088

Post reply on HN