Live data from Hacker News

Furiosa: 3.5x efficiency over H100s

furiosa.ai

131–140 of 165 posts

Re: Furiosa: 3.5x efficiency over H100s

#131
post #7

I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…

> I am of the opinion that Nvidia's hit the wall with their current architecture

Google presented TPUs in 2015. NVIDIA introduced Tensor Cores in 2018. Both utilize systolic arrays.

And last month NVIDIA pseudo-acquired Groq including the founder and original TPU guy. Their LPUs are way more efficient for inference. Also of note Groq is fully made in USA and has a very diverse supply chain using older nodes.

NVIDIA architecture is more than fine. They have deep pockets and very technical leadership. Their weakness lies more with their customers, lack of energy, and their dependency on TSMC and the memory cartel.

Re: Furiosa: 3.5x efficiency over H100s

#132
post #7

I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…

What about Etched? https://techfundingnews.com/nvidia-rival-ai-chip-maker-etche...

Re: Furiosa: 3.5x efficiency over H100s

#133

Earlier quoted context omitted.

The circle allows you to put an arbitrary "price" on those services. You could say that the bathroom and table are $100 each, so your combined work was $200. Or you could claim that each of you did $1M work. Without actual money flowing in/out of your circle, your claims aren't tethered to reality.

You don’t think real money is changing hands when Microsoft buys Nvidia GPUs?

It's a soft version of money printing basically. These firms are clearly inflating each other's valuations by making huge promises of future business to each other. Naively, one would look at the headlines and draw the conclusion that much more money is going to flow into AI in the near future.

Of course, a rational investor looks at this and discounts the fact that most of those promises are predicated on insane growth that has no grounding in reality.

However, there are plenty of greedy or irrational investors, whose recklessness will affect everyone, not just them.

Re: Furiosa: 3.5x efficiency over H100s

#134

Earlier quoted context omitted.

The circle allows you to put an arbitrary "price" on those services. You could say that the bathroom and table are $100 each, so your combined work was $200. Or you could claim that each of you did $1M work. Without actual money flowing in/out of your circle, your claims aren't tethered to reality.

You don’t think real money is changing hands when Microsoft buys Nvidia GPUs?

What about when Nvidia sells GPUs to a client and then buys 10% of their shares?

Re: Furiosa: 3.5x efficiency over H100s

#135
post #45

Earlier quoted context omitted.

5.2 is great if you ask it engineering questions, or questions an engineer might ask. It is extremely mid, and actually worse than the o3/o4 era models if you start asking it trivia like if the I-80 tunnel on the bay bridge (yerba buena island) is the largest bore in the world. Don't even get me started on whatever model is wired up to the voice chat button. But yes it will write you a flawless, physics accurate flig…

My impression is that software developers are the lions share of people actually paying for AI, but perhaps that's just my bubble world view.

According to OpenAI it's something like 4.2% of the use. But this data is from before Codex added subscription support and I think only covers ChatGPT (back when most people were using ChatGPT for coding work, before agents got good).

https://i.imgur.com/0XG2CKE.jpeg

Re: Furiosa: 3.5x efficiency over H100s

#136

These things never pan out. The reasons why this almost never works is one of the following: - They assume they can move hardware complexity (scheduling etc, access patterns into software). The magic compiler/runtime never arrives. - They assume their hard-to-program but faster architecture will get figured out by devs. It won't. - They assume a certain workload. The workload changes, and their arch is no longer opti…

> - They assume their hard-to-program but faster architecture will get figured out by devs. It won't.

Or it will get figured out in the niche fields where people are willing to figure out really hard stuff to squeeze out max performance (PE, hedge funds, intelligence)

Either way agree, it's hard to get mass adoption without the software ecosystem feeding back in

Re: Furiosa: 3.5x efficiency over H100s

#137

Earlier quoted context omitted.

>go download a model GP was talking about commercially hosted LLMs running in datacenters, not free Chinese models. Local is definitely still improving. That’s another reason the megacenter model (NVDA’s big line up forever plan) is either a financial catastrophe about to happen, or the biggest bailout ever.

GPT 5.2 is an incredible leap over 5.1 / 5

how is “GPT 5.2 is good” a response to “downloadable models aren’t relevant”?

Re: Furiosa: 3.5x efficiency over H100s

#138
post #26
post #18

Earlier quoted context omitted.

> I am of the opinion that Nvidia's hit the wall with their current architecture Based on what? Their measured performance on things people care about keep going up, and their software stack keeps getting better and unlocking more performance on existing hardware Inference tests: https://inferencemax.semianalysis.com/ Training tests: https://www.lightly.ai/blog/nvidia-b200-vs-h100 https://newsletter.semianalysis.com/…

> Their measured performance on things people care about keep going up, and their software stack keeps getting better and unlocking more performance on existing hardware I'm more concerned about fully-loaded dollars per token - including datacenter and power costs - rather than "does the chip go faster." If Nvidia couldn't make the chip go faster, there wouldn't be any debate, the question right now is "what is the c…

[deleted]

Re: Furiosa: 3.5x efficiency over H100s

#139
post #26
post #18

Earlier quoted context omitted.

> I am of the opinion that Nvidia's hit the wall with their current architecture Based on what? Their measured performance on things people care about keep going up, and their software stack keeps getting better and unlocking more performance on existing hardware Inference tests: https://inferencemax.semianalysis.com/ Training tests: https://www.lightly.ai/blog/nvidia-b200-vs-h100 https://newsletter.semianalysis.com/…

> Their measured performance on things people care about keep going up, and their software stack keeps getting better and unlocking more performance on existing hardware I'm more concerned about fully-loaded dollars per token - including datacenter and power costs - rather than "does the chip go faster." If Nvidia couldn't make the chip go faster, there wouldn't be any debate, the question right now is "what is the c…

[deleted]

Re: Furiosa: 3.5x efficiency over H100s

#140
post #74

Earlier quoted context omitted.

> nothing about the industry's finances add up right now Nothing about the industry’s finances, or about Anthropic and OpenAI’s finances? I look at the list of providers on OpenRouter for open models, and I don’t believe all of them are losing money. FWIW Anthropic claims (iirc) that they don’t lose money on inference. So I don’t think the industry or the model of selling inference is what’s in trouble there. I am mu…

It's become rather clear from the local LLM communities catching up that there is no moat. Everyone is still just barely figuring out how this nifty data structures produce such a powerful emergent behavior, there isn't any truly secret sauce yet.

> local LLM communities catching up that there is no moat.

they use Chinese open LLMs, but Chinese companies have moat: training datasets and some non-opensource tech, and also salaried talents, which one would need serious investment for if decide to bootstrap competitive frontier model today.

Post reply on HN