Live data from Hacker News

15 years of FP64 segmentation, and why the Blackwell Ultra breaks the pattern

nicolasdickenmann.com

71–80 of 93 posts

Re: 15 years of FP64 segmentation, and why the Blackwell Ultra breaks the pattern

#71

Earlier quoted context omitted.

Nvidia started their GPGPU adventure by acquiring a physics engine and porting it over to run on their GPUs. Supporting linear algebra operations was pretty much the goal from the start.

They were also full of lies when they have started their GPGPU adventure (like also today). For a few years they have repeated continuously how GPGPU can provide about 100 times more speed than CPUs. This has always been false. GPUs are really much faster, but their performance per watt has oscillated during most of the time around 3 times and sometimes up to 4 times greater in comparison with CPUs. This is impressiv…

> GPGPU can provide about 100 times more speed than CPUs

Ok. You're talking about performance.

> their performance per watt has oscillated during most of the time around 3 times and sometimes up to 4 times greater in comparison with CPUs

Now you're talking about perf/W.

> This is impressive, but very far from the "100" factor originally claimed by NVIDIA.

That's because you're comparing apples to apples per apple cart.

Re: 15 years of FP64 segmentation, and why the Blackwell Ultra breaks the pattern

#72
post #65
post #45

Earlier quoted context omitted.

> subsidizing the pro parts. You got this wrong way around. It's the high margin (pro) products subsidizing low margin (consumer) products.

In general, yes, but when consumer parts are spending silicon area on features they can't use, it is happening in the other direction too.

[deleted]

Re: 15 years of FP64 segmentation, and why the Blackwell Ultra breaks the pattern

#73
post #32
post #8

I'm not sure why the article dismisses cost. Let's say X=10% of the GPU area (~75mm^2) is dedicated to FP32 SIMD units. Assume FP64 units are ~2-4x bigger. That would be 150-300mm^2, a huge amount of area that would increase the price per GPU. You may not agree with these assumptions. Feel free to change them. It is an overhead that is replicated per core. Why would gamers want to pay for any features they don't use?…

> Assume FP64 units are ~2-4x bigger. This is wrong assumption. FP64 usually uses the same circuitry as two FP32, adding not that much ((de)normalization, mostly). From the top of my head, overhead is around 10% or so. > Why would gamers want to pay for any features they don't use? https://www.youtube.com/watch?v=lEBQveBCtKY Apparently FP80, which is even wider than FP64, is beneficial for pathfinding algorithms in g…

Has FP80 ever existed anywhere other than x87?

Re: 15 years of FP64 segmentation, and why the Blackwell Ultra breaks the pattern

#74

A question that has been bugging me for a while is what will NVIDIA do with its HPC business? By HPC I mean clusters intended for non-AI related workloads. Are they going to cater to them separetely, or are they going to tell them to just emulate FP64?

For a long time AMD has been offering much better FP64 performance than NVIDIA, in their CDNA GPUs (which continue the older AMD GCN ISA, instead of resembling the RDNA used in gaming GPUs).

Nevertheless, the AMD GPUs continue to have their old problems, weak software support, so-and-so documentation, software incompatibility with the cheap GPUs that could be used directly by a programmer for developing applications.

There is a promise that AMD will eventually unify the ISA of their "datacenter" and gaming GPUs, like NVIDIA has always done, but it in unclear when this will happen.

Thus they are a solution only for big companies or government agencies.

Re: 15 years of FP64 segmentation, and why the Blackwell Ultra breaks the pattern

#75

A question that has been bugging me for a while is what will NVIDIA do with its HPC business? By HPC I mean clusters intended for non-AI related workloads. Are they going to cater to them separetely, or are they going to tell them to just emulate FP64?

AMD MI430X is taking that market.

Re: 15 years of FP64 segmentation, and why the Blackwell Ultra breaks the pattern

#76
post #66
post #2

It's amazing to step back and look at how much of NVIDIA's success has come from unforeseen directions. For their original purpose of making graphics chips, the consumer vs pro divide was all about CAD support and optional OpenGL features that games didn't use. Programmable shaders were added for the sake of graphics rendering needs, but ended up spawning the whole GPGPU concept, which NVIDIA reacted to very well wit…

The counter question is: why have AMD been so bad by comparison?

Because GPUs require a lot on the software side, and AMD sucks at software. They are a CPU company that bought a GPU company. ATI should have been left alone.

Re: 15 years of FP64 segmentation, and why the Blackwell Ultra breaks the pattern

#77
post #65
post #45

Earlier quoted context omitted.

> subsidizing the pro parts. You got this wrong way around. It's the high margin (pro) products subsidizing low margin (consumer) products.

In general, yes, but when consumer parts are spending silicon area on features they can't use, it is happening in the other direction too.

[deleted]

Re: 15 years of FP64 segmentation, and why the Blackwell Ultra breaks the pattern

#78
post #3

FP64 performance is limited on consumer because the US government deems it important to nuclear weapons research. Past a certain threshold of FP64 throughput, your chip goes in a separate category and is subject to more regulation about who you can sell to and know-your-customer. FP32 does not matter for this threshold. https://en.wikipedia.org/wiki/Adjusted_Peak_Performance It is not a market segmentation tactic and…

It's surprising that this restriction continues to linger at all. The newest nuclear warhead models in the US arsenal were developed in the 1970s, when supercomputer performance was well below 1 gigaflop. When the US stopped testing nuclear warheads in 1992, top end supercomputers were under 10 gigaflops. The only thing the US arsenal needs faster computers for is simulating the behavior of its aging warhead stockpile without physical tests, which is not going to matter to a state building its first nuclear weapons.

Re: 15 years of FP64 segmentation, and why the Blackwell Ultra breaks the pattern

#79
post #17
post #9

Earlier quoted context omitted.

Most people don't appreciate how many dead end applications NVIDIA explored before finding deep learning. It took a very long time, and it wasn't luck.

It was luck that a viable non-graphics application like deep learning existed which was well-suited to the architecture NVIDIA already had on hand. I certainly don't mean to diminish the work NVIDIA did to build their CUDA ecosystem, but without the benefit of hindsight I think it would have been very plausible that GPU architectures would not have been amenable to any use cases that would end up dwarfing graphics it…

There's something of a feedback loop here, in that the reason that transformers and attention won over all the other forms of AI/ML is that they worked very well on the architecture that NVIDIA had already built, so you could scale your model size very dramatically just by throwing more commodity hardware at it.

Re: 15 years of FP64 segmentation, and why the Blackwell Ultra breaks the pattern

#80
post #65
post #45

Earlier quoted context omitted.

> subsidizing the pro parts. You got this wrong way around. It's the high margin (pro) products subsidizing low margin (consumer) products.

In general, yes, but when consumer parts are spending silicon area on features they can't use, it is happening in the other direction too.

This isn't really true, and it wouldn't be a big deal even if it was.

Die areas for consumer card chips are smaller than die areas for datacenter card chips, and this has held for a few generations now. They can't possibly be the same chips, because they are physically different sizes. The lowest-end consumer dies are less than 1/4 the area of datacenter dies, and even the highest-end consumer dies are only like 80% the area of datacenter dies. This implies there must be some nontrivial differentiation going on at the silicon level.

Secondly, you are not paying for the die area anyway. Whether a chip is obtained from being specially made for that exact model of GPU, or it is obtained from being binned after possibly defective areas get fused off, you are paying for the end-result product. If that product meets the expected performance, it is doing its job. This is not a subsidy (at least, not in that direction), the die is just one small part of what makes a usable GPU card, and excess die area left dark isn't even pure waste, as it helps with heat dissipation.

The fact that nVidia excludes decent FP64 from all of its prosumer offerings (*) can still be called "artificial" insofar as it was indeed done on purpose for market segmentation purposes, but it's not some trivial trick. They really are just not putting it into the silicon. This has been the case for longer than it wasn't by now, even.

* = The Quadro line of "professional" workstation cards nowadays are just consumer cards with ECC RAM and special drivers

Post reply on HN