Earlier quoted context omitted.
Inference costs scale linearly with usage. R&D expenses do not. That's not to mention that Dario Amodei has said that their models actually have a good return, even when accounting for training costs [0]. [0] https://youtu.be/GcqQ1ebBqkc?si=Vs2R4taIhj3uwIyj&t=1088
> Inference costs scale linearly with usage. R&D expenses do not. Do we know this is true for AI?
Furiosa: 3.5x efficiency over H100s
111–120 of 165 posts
Re: Furiosa: 3.5x efficiency over H100s
#112I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…
Based on conversations I've had with some people managing GPU's at scale in the datacenters, inference is an after thought. There is a gold rush for training right now, and that's where these massive clusters are being used. LLM's are probably a small fraction of the overall GPU compute in use right now. I suspect in the next 5 years we'll have full Hollywood movies being completely generated (at least the specialfx)…
Re: Furiosa: 3.5x efficiency over H100s
#113Earlier quoted context omitted.
Hollywood studios are breathing their last gasps now. Anyone will be able to use AI to create blockbuster type movies, Hollywood's moat around that is rapidly draining.
Anybody had the ability to write the next great novel for a while, but few succeed.
The other thing to compare is the narrative quality. I find even middling books to be of much higher quality than blockbuster movies on average. Or rather I'm constantly appalled at what passes for a decent script. I assume that's due to needing to appeal to a broad swath of the population because production is so expensive, but understanding the (likely) reason behind it doesn't do anything to improve the end result.
So if "all" we get out of this is a 1000x reduction in production budgets which leads to a 100x increase in the amount of media available I expect it will be a huge win for the consumer.
Re: Furiosa: 3.5x efficiency over H100s
#114Earlier quoted context omitted.
> Their measured performance on things people care about keep going up, and their software stack keeps getting better and unlocking more performance on existing hardware I'm more concerned about fully-loaded dollars per token - including datacenter and power costs - rather than "does the chip go faster." If Nvidia couldn't make the chip go faster, there wouldn't be any debate, the question right now is "what is the c…
> OpenAI has $1.15T in spend commitments over the next 10 years Yes, but those aren't contracted commitments, and we know some of them are equity swaps. For example "Microsoft ($250B Azure commitment)" from the footnote is an unknown amount of actual cash. And I think it's fair to point out the other information in your link "OpenAI projects a 48% gross profit margin in 2025, improving to 70% by 2029."
The economics of the entire setup are laughable and it's obvious that it's a massive bubble. The profit that'd need to be delivered to justify the current valuations is far beyond what is actually realistic.
What moat does OpenAI have? I'd argue basically none. They make extremely lofty forecasts and project an image of crazy growth opportunities, but is that going to ever survive the bubble popping?
Re: Furiosa: 3.5x efficiency over H100s
#115Earlier quoted context omitted.
> But when you just look at it from an inference perspective, looking at these data centres like token factories makes sense. So if you ignore the majority of the costs, then it makes sense. Opus 4.5 was released on November 25, 2025. That is less than 2 months ago. When they stop training new models, then we can forget about training costs.
I'm not taking a side here - I don't know enough - but it's an interesting line of reasoning. So I'll ask, how is that any different than fabs? From what I understand R&D is absurd and upgrading to a new node is even more absurd. The resulting chips sell for chump change on a per unit basis (analogous to tokens). But somehow it all works out. Well, sort of. The bleeding edge companies kept dropping out until you coul…
Invariably, there's going to be a collapse in the hype, the bubble will burst, and an investment deleveraging will remove a lot of money from the space in a short period of time. The bigger the bubble, the more painful and less survivable this event will be.
Re: Furiosa: 3.5x efficiency over H100s
#116Earlier quoted context omitted.
> OpenAI has $1.15T in spend commitments over the next 10 years Yes, but those aren't contracted commitments, and we know some of them are equity swaps. For example "Microsoft ($250B Azure commitment)" from the footnote is an unknown amount of actual cash. And I think it's fair to point out the other information in your link "OpenAI projects a 48% gross profit margin in 2025, improving to 70% by 2029."
The fact that there's an incestual circle between OpenAI, Microsoft, NVidia, AMD, etc.. where they provide massive promises to each other for future business is nothing short of hilarious. The economics of the entire setup are laughable and it's obvious that it's a massive bubble. The profit that'd need to be delivered to justify the current valuations is far beyond what is actually realistic. What moat does OpenAI h…
Re: Furiosa: 3.5x efficiency over H100s
#117Earlier quoted context omitted.
The fact that there's an incestual circle between OpenAI, Microsoft, NVidia, AMD, etc.. where they provide massive promises to each other for future business is nothing short of hilarious. The economics of the entire setup are laughable and it's obvious that it's a massive bubble. The profit that'd need to be delivered to justify the current valuations is far beyond what is actually realistic. What moat does OpenAI h…
I still don't really understand this "circle" issue. If I fix your bathroom and in return you make me a new table, is that an incestuous circle? Haven't we both just exchanged value?
Re: Furiosa: 3.5x efficiency over H100s
#118Earlier quoted context omitted.
> Inference costs scale linearly with usage. R&D expenses do not. Do we know this is true for AI?
Yes. R&D is guaranteed to fall as a percentage of costs eventually. The only question is when, and there is also a question of who is still solvent when that time comes. It is competition and an innovation race that keeps it so high, and it won't stay so high forever. Either rising revenues or falling competition will bring R&D costs down as a percentage of revenue at some point.
Re: Furiosa: 3.5x efficiency over H100s
#119Earlier quoted context omitted.
> (at which point Intel typically fired their chip design group, hired everyone from AMD or whoever, and came out with Core or whatever) Didn't the Core architecture come from the Intel Pentium M Israeli team? https://en.wikipedia.org/wiki/Intel_Core_(microarchitecture)...
Yeah, that bit was pure snark - point was Intel’s gotten caught resting on their laurels a couple times when their architectures get a little long in the tooth, and often it’s existential enough that the team that pulls them out of it isn’t the one that put them in it.
If you wanted to make that point, Itanium or 64-bit/multi-core desktop processing would be better examples than Core.
Re: Furiosa: 3.5x efficiency over H100s
#120Earlier quoted context omitted.
> OpenAI has $1.15T in spend commitments over the next 10 years Yes, but those aren't contracted commitments, and we know some of them are equity swaps. For example "Microsoft ($250B Azure commitment)" from the footnote is an unknown amount of actual cash. And I think it's fair to point out the other information in your link "OpenAI projects a 48% gross profit margin in 2025, improving to 70% by 2029."
> "OpenAI projects a 48% gross profit margin in 2025, improving to 70% by 2029." OpenAI can project whatever they want, they're not public.
Private companies do have a license to lie to their shareholders.