Live data from Hacker News

Furiosa: 3.5x efficiency over H100s

furiosa.ai

111–120 of 165 posts

Re: Furiosa: 3.5x efficiency over H100s

#111

Earlier quoted context omitted.

Inference costs scale linearly with usage. R&D expenses do not. That's not to mention that Dario Amodei has said that their models actually have a good return, even when accounting for training costs [0]. [0] https://youtu.be/GcqQ1ebBqkc?si=Vs2R4taIhj3uwIyj&t=1088

> Inference costs scale linearly with usage. R&D expenses do not. Do we know this is true for AI?

Yes. R&D is guaranteed to fall as a percentage of costs eventually. The only question is when, and there is also a question of who is still solvent when that time comes. It is competition and an innovation race that keeps it so high, and it won't stay so high forever. Either rising revenues or falling competition will bring R&D costs down as a percentage of revenue at some point.

Re: Furiosa: 3.5x efficiency over H100s

#112
post #7

I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…

Based on conversations I've had with some people managing GPU's at scale in the datacenters, inference is an after thought. There is a gold rush for training right now, and that's where these massive clusters are being used. LLM's are probably a small fraction of the overall GPU compute in use right now. I suspect in the next 5 years we'll have full Hollywood movies being completely generated (at least the specialfx)…

it's so weird how they spend all this money to train new models and then open sources it. it's gold rush but nvidia is getting all the gold.

Re: Furiosa: 3.5x efficiency over H100s

#113
post #70

Earlier quoted context omitted.

Hollywood studios are breathing their last gasps now. Anyone will be able to use AI to create blockbuster type movies, Hollywood's moat around that is rapidly draining.

Anybody had the ability to write the next great novel for a while, but few succeed.

There are lots of very good relatively recent novels on the shelf at the bookstore. Certainly orders of magnitude more than there are movies.

The other thing to compare is the narrative quality. I find even middling books to be of much higher quality than blockbuster movies on average. Or rather I'm constantly appalled at what passes for a decent script. I assume that's due to needing to appeal to a broad swath of the population because production is so expensive, but understanding the (likely) reason behind it doesn't do anything to improve the end result.

So if "all" we get out of this is a 1000x reduction in production budgets which leads to a 100x increase in the amount of media available I expect it will be a huge win for the consumer.

Re: Furiosa: 3.5x efficiency over H100s

#114
post #33
post #26

Earlier quoted context omitted.

> Their measured performance on things people care about keep going up, and their software stack keeps getting better and unlocking more performance on existing hardware I'm more concerned about fully-loaded dollars per token - including datacenter and power costs - rather than "does the chip go faster." If Nvidia couldn't make the chip go faster, there wouldn't be any debate, the question right now is "what is the c…

> OpenAI has $1.15T in spend commitments over the next 10 years Yes, but those aren't contracted commitments, and we know some of them are equity swaps. For example "Microsoft ($250B Azure commitment)" from the footnote is an unknown amount of actual cash. And I think it's fair to point out the other information in your link "OpenAI projects a 48% gross profit margin in 2025, improving to 70% by 2029."

The fact that there's an incestual circle between OpenAI, Microsoft, NVidia, AMD, etc.. where they provide massive promises to each other for future business is nothing short of hilarious.

The economics of the entire setup are laughable and it's obvious that it's a massive bubble. The profit that'd need to be delivered to justify the current valuations is far beyond what is actually realistic.

What moat does OpenAI have? I'd argue basically none. They make extremely lofty forecasts and project an image of crazy growth opportunities, but is that going to ever survive the bubble popping?

Re: Furiosa: 3.5x efficiency over H100s

#115

Earlier quoted context omitted.

> But when you just look at it from an inference perspective, looking at these data centres like token factories makes sense. So if you ignore the majority of the costs, then it makes sense. Opus 4.5 was released on November 25, 2025. That is less than 2 months ago. When they stop training new models, then we can forget about training costs.

I'm not taking a side here - I don't know enough - but it's an interesting line of reasoning. So I'll ask, how is that any different than fabs? From what I understand R&D is absurd and upgrading to a new node is even more absurd. The resulting chips sell for chump change on a per unit basis (analogous to tokens). But somehow it all works out. Well, sort of. The bleeding edge companies kept dropping out until you coul…

Someone else mentioned it elsewhere in this thread, and I believe this is the crux of the issue: this is all predicated in the actual end users finding enough benefit in LLM services to keep the gravy train going. It's irrelevant how scalable and profitable the shovel makes are, to keep this business afloat long term, the shovelers - ie the end users - have to make money using the shovesl. Those expectations are currently ridiculously inflated. Far beyond anything in the past.

Invariably, there's going to be a collapse in the hype, the bubble will burst, and an investment deleveraging will remove a lot of money from the space in a short period of time. The bigger the bubble, the more painful and less survivable this event will be.

Re: Furiosa: 3.5x efficiency over H100s

#116
post #33

Earlier quoted context omitted.

> OpenAI has $1.15T in spend commitments over the next 10 years Yes, but those aren't contracted commitments, and we know some of them are equity swaps. For example "Microsoft ($250B Azure commitment)" from the footnote is an unknown amount of actual cash. And I think it's fair to point out the other information in your link "OpenAI projects a 48% gross profit margin in 2025, improving to 70% by 2029."

The fact that there's an incestual circle between OpenAI, Microsoft, NVidia, AMD, etc.. where they provide massive promises to each other for future business is nothing short of hilarious. The economics of the entire setup are laughable and it's obvious that it's a massive bubble. The profit that'd need to be delivered to justify the current valuations is far beyond what is actually realistic. What moat does OpenAI h…

I still don't really understand this "circle" issue. If I fix your bathroom and in return you make me a new table, is that an incestuous circle? Haven't we both just exchanged value?

Re: Furiosa: 3.5x efficiency over H100s

#117

Earlier quoted context omitted.

The fact that there's an incestual circle between OpenAI, Microsoft, NVidia, AMD, etc.. where they provide massive promises to each other for future business is nothing short of hilarious. The economics of the entire setup are laughable and it's obvious that it's a massive bubble. The profit that'd need to be delivered to justify the current valuations is far beyond what is actually realistic. What moat does OpenAI h…

I still don't really understand this "circle" issue. If I fix your bathroom and in return you make me a new table, is that an incestuous circle? Haven't we both just exchanged value?

The circle allows you to put an arbitrary "price" on those services. You could say that the bathroom and table are $100 each, so your combined work was $200. Or you could claim that each of you did $1M work. Without actual money flowing in/out of your circle, your claims aren't tethered to reality.

Re: Furiosa: 3.5x efficiency over H100s

#118

Earlier quoted context omitted.

> Inference costs scale linearly with usage. R&D expenses do not. Do we know this is true for AI?

Yes. R&D is guaranteed to fall as a percentage of costs eventually. The only question is when, and there is also a question of who is still solvent when that time comes. It is competition and an innovation race that keeps it so high, and it won't stay so high forever. Either rising revenues or falling competition will bring R&D costs down as a percentage of revenue at some point.

Yes, but eventually may be longer than the market can hold out. So far R&D expenses have skyrocketed and it does not look like that will be changing anytime soon.

Re: Furiosa: 3.5x efficiency over H100s

#119
post #64
post #55

Earlier quoted context omitted.

> (at which point Intel typically fired their chip design group, hired everyone from AMD or whoever, and came out with Core or whatever) Didn't the Core architecture come from the Intel Pentium M Israeli team? https://en.wikipedia.org/wiki/Intel_Core_(microarchitecture)...

Yeah, that bit was pure snark - point was Intel’s gotten caught resting on their laurels a couple times when their architectures get a little long in the tooth, and often it’s existential enough that the team that pulls them out of it isn’t the one that put them in it.

I think that's an overly reductive view of a very complicated problem space, with the benefit of hindsight.

If you wanted to make that point, Itanium or 64-bit/multi-core desktop processing would be better examples than Core.

Re: Furiosa: 3.5x efficiency over H100s

#120
post #42
post #33

Earlier quoted context omitted.

> OpenAI has $1.15T in spend commitments over the next 10 years Yes, but those aren't contracted commitments, and we know some of them are equity swaps. For example "Microsoft ($250B Azure commitment)" from the footnote is an unknown amount of actual cash. And I think it's fair to point out the other information in your link "OpenAI projects a 48% gross profit margin in 2025, improving to 70% by 2029."

> "OpenAI projects a 48% gross profit margin in 2025, improving to 70% by 2029." OpenAI can project whatever they want, they're not public.

They still have shareholders who can sue for misinformation.

Private companies do have a license to lie to their shareholders.

Post reply on HN