Earlier quoted context omitted.
GPT 5.2 is an incredible leap over 5.1 / 5
5.2 is great if you ask it engineering questions, or questions an engineer might ask. It is extremely mid, and actually worse than the o3/o4 era models if you start asking it trivia like if the I-80 tunnel on the bay bridge (yerba buena island) is the largest bore in the world. Don't even get me started on whatever model is wired up to the voice chat button. But yes it will write you a flawless, physics accurate flig…
Furiosa: 3.5x efficiency over H100s
51–60 of 165 posts
Re: Furiosa: 3.5x efficiency over H100s
#52Earlier quoted context omitted.
GPT 5.2 is an incredible leap over 5.1 / 5
5.2 is great if you ask it engineering questions, or questions an engineer might ask. It is extremely mid, and actually worse than the o3/o4 era models if you start asking it trivia like if the I-80 tunnel on the bay bridge (yerba buena island) is the largest bore in the world. Don't even get me started on whatever model is wired up to the voice chat button. But yes it will write you a flawless, physics accurate flig…
Re: Furiosa: 3.5x efficiency over H100s
#53I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…
Based on conversations I've had with some people managing GPU's at scale in the datacenters, inference is an after thought. There is a gold rush for training right now, and that's where these massive clusters are being used. LLM's are probably a small fraction of the overall GPU compute in use right now. I suspect in the next 5 years we'll have full Hollywood movies being completely generated (at least the specialfx)…
Re: Furiosa: 3.5x efficiency over H100s
#54I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…
We've seen this before. In 2001, there were something like 50+ OC-768 hardware startups. At the time, something like 5 OC-768 links could carry all the traffic in the world . Even exponential doubling every 12 months wasn't going to get enough customers to warrant all the funding that had poured into those startups. When your business model bumps into "All the in the world," you're in trouble.
Re: Furiosa: 3.5x efficiency over H100s
#55I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…
Didn't the Core architecture come from the Intel Pentium M Israeli team? https://en.wikipedia.org/wiki/Intel_Core_(microarchitecture)...
Re: Furiosa: 3.5x efficiency over H100s
#56really weird graph where they're comparing to 3x H100 PCI-E which is a config I don't think anyone is using. they're trying to compare at iso-power? I just want to see their box vs a box of 8 h100s b/c that's what people would buy instead, and they can divide tokens and watts if that's the pitch.
Yeah they are defining a "rack" as 15kW, though 3x H100 PCIe is only a bit over 1kW. So they are assuming GPUs are <10% of rack power usage which sounds suspiciously low.
Re: Furiosa: 3.5x efficiency over H100s
#57I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…
> (at which point Intel typically fired their chip design group, hired everyone from AMD or whoever, and came out with Core or whatever) Didn't the Core architecture come from the Intel Pentium M Israeli team? https://en.wikipedia.org/wiki/Intel_Core_(microarchitecture)...
Re: Furiosa: 3.5x efficiency over H100s
#58This is from September 2025, what's new?
Hence the relevance, maybe.
Re: Furiosa: 3.5x efficiency over H100s
#59Earlier quoted context omitted.
That may not be a good example because everyone is saying Groq isn't worth $20B.
They were valued at $6.9B just three months before Nvidia bought them for $20B, triple the valuation. That figure seems to have been pulled out of thin air.
Re: Furiosa: 3.5x efficiency over H100s
#60Earlier quoted context omitted.
Based on conversations I've had with some people managing GPU's at scale in the datacenters, inference is an after thought. There is a gold rush for training right now, and that's where these massive clusters are being used. LLM's are probably a small fraction of the overall GPU compute in use right now. I suspect in the next 5 years we'll have full Hollywood movies being completely generated (at least the specialfx)…
Hollywood studios are breathing their last gasps now. Anyone will be able to use AI to create blockbuster type movies, Hollywood's moat around that is rapidly draining.