Live data from Hacker News

Furiosa: 3.5x efficiency over H100s

furiosa.ai

151–160 of 165 posts

Re: Furiosa: 3.5x efficiency over H100s

#151

Earlier quoted context omitted.

Consensus seems to be that the labs are profitable on inference. They are only losing money on training and free users. The competition requiring them to spend that money on training and free users does complicate things. But when you just look at it from an inference perspective, looking at these data centres like token factories makes sense. I would definitely pay more to get faster inference of Opus 4.5, for examp…

>Consensus seems to be that the labs are profitable on inference. They are only losing money on training and free users. That sounds like “we’re profitable if you ignore our biggest expenses.” If they could be profitable now, we’d see at least a few companies just be profitable and stop the heavy expenses. My guess is it’s simply not the case or everyone’s trapped in a cycle where they are all required to keep spendi…

This is just not true. Plenty of companies will remain unprofitable for as long as they can in the name of growth, market share, and beating their competition. At some point it will level out, but while they can still raise cheap capital and spend it to grow, they will.

OpenAI could put in ads tomorrow and make tons of money overnight. The only reason they don't is competition. But when they start to find it harder to raise capital to fund their growth, they will.

Re: Furiosa: 3.5x efficiency over H100s

#152

Earlier quoted context omitted.

Yes. R&D is guaranteed to fall as a percentage of costs eventually. The only question is when, and there is also a question of who is still solvent when that time comes. It is competition and an innovation race that keeps it so high, and it won't stay so high forever. Either rising revenues or falling competition will bring R&D costs down as a percentage of revenue at some point.

Yes, but eventually may be longer than the market can hold out. So far R&D expenses have skyrocketed and it does not look like that will be changing anytime soon.

That's why it is a bet, and not a sure thing.

Re: Furiosa: 3.5x efficiency over H100s

#153
post #45

Earlier quoted context omitted.

5.2 is great if you ask it engineering questions, or questions an engineer might ask. It is extremely mid, and actually worse than the o3/o4 era models if you start asking it trivia like if the I-80 tunnel on the bay bridge (yerba buena island) is the largest bore in the world. Don't even get me started on whatever model is wired up to the voice chat button. But yes it will write you a flawless, physics accurate flig…

But how many are willing to fork over $20 or so a month to ask simple trivia questions?

In addition to engineering tasks, it's an ad-free answer-box, outside of cross checking things, or browsing search results it's totally replaced Google/search engine use for me. I also pay for Kagi for search. In the last year I've been able to fully divorce myself from the google ecosystem besides gmail and maps.

Re: Furiosa: 3.5x efficiency over H100s

#154

Earlier quoted context omitted.

My impression is that software developers are the lions share of people actually paying for AI, but perhaps that's just my bubble world view.

According to OpenAI it's something like 4.2% of the use. But this data is from before Codex added subscription support and I think only covers ChatGPT (back when most people were using ChatGPT for coding work, before agents got good). https://i.imgur.com/0XG2CKE.jpeg

The execs I've talked to, they are paying for it to answer capex questions, as a sounding board for decision making, and perhaps most importantly, crafting/modifying emails for tone/content. In the bay area particularly a lot of execs are foreign with english as their second language and LLMs can cut email generation time in half.

Re: Furiosa: 3.5x efficiency over H100s

#155
post #78

Earlier quoted context omitted.

> Nothing they create make any goddamn sense, I wouldn’t be that dismissive. Some have managed to make impressive things with them (although nothing close to an actual movie, even a short). https://www.youtube.com/watch?v=ET7Y1nNMXmA A bit older: https://www.youtube.com/watch?v=8OOpYvxKhtY Compared to two years ago: https://www.youtube.com/watch?v=LHeCTfQOQcs

The problem with all of these, even the most recent one, is that they have the "AI look". People have tired of this look already, even for short adverts; if they don't want five minutes of it, they really won't like two hours of it. There is no doubt the quality has vastly improved over time, but I see no sign of progress in removing the "AI look" from these things.

My feeling is the definition of the "AI look" has evolved as these models progressed.

It used to mean psychedelic weird things worthy of the strangest dreams or an acid trip.

Then it meant strangely blurry with warped alien script and fifteen fingers, including one coming out of another’s second phalanx

Now it means something odd, off, somewhat both hard to place and obvious, like the CGI "transparent" car (is it that the 3D model is too simple, looks like a bad glass sculpture, and refracts light in squares?) and ice cliffs (I think the the lighting is completely off, and the colours are wrong) in Die Another Day.

And if that’s the case, then these models have covered far more in far less time then it took computer graphics and CGI.

Re: Furiosa: 3.5x efficiency over H100s

#156

Earlier quoted context omitted.

Whats a more realistic config?

8xGPUs per box. this has been the data center standard for the last 8ish years. furthermore usually NVLink connected within the box (SXM instead of PCIe cards, although the physical data link is still PCIe.) this is important because the daughter board provides PCIe switches which usually connect NVMe drives, NICs and GPUs together such that within that subcomplex there isn't any PCIe oversubscription. since last yea…

Fascinating! So each GPU is partnered with disk and NICs such that theres no oversubscription for bandwidth within its 'slice'? (idk what the word is) And each of these 8 slices wire up to NVLink back to the host?

Feels like theres some amount of (software) orchestration for making data sit on the right drives or traverse the right NICs, guess I never really thought about the complexity of this kind of scale.

I googled GB200, its cool that Nvidia sells you a unit rather than expecting you to DIY PC yourself.

Re: Furiosa: 3.5x efficiency over H100s

#157

Earlier quoted context omitted.

According to OpenAI it's something like 4.2% of the use. But this data is from before Codex added subscription support and I think only covers ChatGPT (back when most people were using ChatGPT for coding work, before agents got good). https://i.imgur.com/0XG2CKE.jpeg

I'd believe that but I was commenting on who actually pays for it. My guess is that most individuals using AI in their personal lives are using some sort of free tier.

Yes 95% are unpaid

Re: Furiosa: 3.5x efficiency over H100s

#158
post #134

Earlier quoted context omitted.

What about when Nvidia sells GPUs to a client and then buys 10% of their shares?

Their shares will be based on the client's valuation, which in public markets is externally priced. If not in public markets it is murkier, but will be grounded in some sort of reality so Nvidia gets the right amount of the company.

My point was that's an indirect subsidy. NVIDIA is selling at a discount to prop up their clients.

Re: Furiosa: 3.5x efficiency over H100s

#159
Are all these improvement over custom kernel efficiency code ? Can we bring these to consumer RTX and Pro cards ?

After I read the article :) The improvements in FuriosaAI's NXT RNGD Server are primarily driven by hardware innovations, not software or code changes.

Re: Furiosa: 3.5x efficiency over H100s

#160
post #7

I am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research divis…

What about TPUs? They are more efficient than nvidia GPUs, a huge amount of inference is done with them, and while they are not literally being sold to the public, the whole technology should be influencing the next steps of Nvidia just like AMD influenced Intel

I believe Furiosa hardware is a variant of TPU, and that's pretty much why they can beat Nvidia on LLM inference. GPUs are more general purpose, which means Furiosa can cut some of the overhead with designs dedicated for matrix multiplication only.

https://furiosa.ai/blog/tensor-contraction-processor-ai-chip...

Post reply on HN