Earlier quoted context omitted.
It’s because CPUs tend to be fundamentally limited in the ways they can efficiently utilize transistors to scale performance, by things like cache latency, reorder depth, branch prediction, etc. While gpus have always been god’s strongest soldier for putting transistors on silicon scalably. They were the perfect machine for a world where transistors per dollar doubled every 18 months. On the other hand now that it’s…
So if they're running into hardware fab limits, how does running deep learning on that limited hardware equate to a doubling of price? I don't quite follow the logic there? Yeah the upscaling is nice, but I dunno about $1200 nice... It was cool when they instead focused on things like creating mobile smaller versions of the cards with laptop level power draws, or better cooling systems that weren't as noisy etc. Whil…
Because moore's law wasn't just about transistor count but about the economic impact of exponential growth in transistors-per-$. In a world without moore's law, using more transistors will result in a higher-cost product. If you want to hold product cost fixed, or even contain the cost spiral, you need to do more with the same amount of transistors - performance-per-transistor is the metric that matters now.
AMD and NVIDIA have already stripped down their pure-raster implementation as far as they can go, with RDNA1 and Maxwell respectively. Maxwell actually cut too far (software scheduling, "minimal" DX12 support, etc) honestly. So where do you keep making perf/tr gains after that?
The gaming world has already pretty well settled on TAA (although some people will never accept it) and upscaling is already common in the console world. So, do TAA upscaling better such that you get the performance gains but not the reduction in visual quality that usually comes with it.
Tensor makes up a relatively small amount of die area (5.9% of total Turing die area, based on comparisons between Turing Major/RTX and Turing Minor/GTX SM engine die shots). And that gets you to about 30% faster than FSR2 for a given level of visual output quality. So the perf-per-transistor metric increases. Also, unlike a fixed-function accelerator, it can be used for all kinds of other stuff too. It's basically a whole programmable sub-processor, an accelerator for your accelerator.
https://www.reddit.com/r/hardware/comments/baajes/rtx_adds_1...
Now, why ML as opposed to just running it on shaders? Same logic as adding an AVX unit, math density is a lot higher and it can do a lot of work for applications that are specifically tailored to it. DLSS2 uses a relatively standard TAAU (similar to FSR2) but determines the weighting of the samples using a neural net. This produces a lot higher quality than a procedural algorithm currently can - especially under "bad conditions" like higher degrees of upscaling, low framerate/limited sample count, or temporally unstable/high-temporal-frequency areas of the image.
http://behindthepixels.io/assets/files/DLSS2.0.pdf
https://raw.githubusercontent.com/NVIDIA/DLSS/main/doc/DLSS_...
FSR2 does ok at 4K quality mode, but at 1440p and (especially) 1080p output resolutions and in performance modes it does much worse. FSR2 quality 1080p is more like DLSS2 performance mode or maybe balanced mode, so NVIDIA gets more speedup at a given level of visual quality. And DLAA can produce a better-than-native image when running with a native input quality.
The neural weighting just is a lot more efficient at using its samples, it understands what is going on in the scene (moving edges/occlusion etc) and can extract a higher signal-to-noise ratio from the samples and the textures. It's like an op-amp, the ratio of input:output pixels is the "gain factor", and FSR2 and other traditional TAAU algorithms are simply noisier at any given level of gain, whether that's unity or extreme gain, and have other edge-cases like turn-on threshold (bad performance with low samples). ML is the "schottky diode" of graphics amplification (dangerously mixed metaphor, lol), it's simply a lot more agile at shaping the signal than what came before.
(and while on paper plenty of people have argued that procedural programs should be able to do anything ML can, it's not like AMD and others haven't tried to improve TAAU with FSR2, and many others before them. DLSS2 is better, just like LLMs and Stable Diffusion are a lot better than procedural algorithms in their own niches.)
--
All of this exists completely orthogonally to actual wafer costs or packaging or other things. Packaging may boost that transistors/$ metric a little bit but that just gives you a little more to play with. It does allow you to make chips with twice the transistors at twice the cost and have them yield at high rates, but fundamentally 2x400mm2 is still 800mm2 of silicon even if you yield at 100% - you're using more wafer, which drives up costs. Wafer costs have been increasing nearly as fast as density (and predicted to match/pass at 3nm) but there has been a small gain in tr/$, certainly nowhere near the rate of moore's law days. But if wafers cost 8x what they did for 28nm, and are continuing to increase at ~50% per generation, and you keep using more wafer area to compensate for slowing shrinks, then costs will go up (even more than they have).
There is no direct link between "running deep learning" and costs going up. That is just happening independently and affects AMD too, even when they didn't go in on deep learning. TSMC prices keep going up (as do their margins, even now) and even if they went to zero, the design+validation costs are still going up too. A lot of these costs are driven by hard physics problems and not just TSMC profit margin (although it doesn't help).
--
> It was cool when they instead focused on things like creating mobile smaller versions of the cards with laptop level power draws, or better cooling systems that weren't as noisy etc.
I think a 30% performance boost at native visual quality, without power increase is pretty cool. Don't laptops benefit from having 30% higher perf/w just from turning on a setting? And Ada itself is a ~60% perf/w increase over previous generations too.
Like Ada is one of the most efficiency-focused generations ever, much moreso than Ampere or Turing with their trailing nodes. DLSS just stacks on top of this - and unlike FSR2, NVIDIA doesn't fall apart at 1080p resolutions that laptops tend to be using.
Cost is higher than people want, but on the other hand (a) that's going to be the reality unless there is a breakthrough in transistors-per-$, you can't make a fixed number of transistors infinitely fast, there is some asymptotic limit. And (b) people are cherrypicking favored examples or comparing against trailing-node products that had larger dies on slower, less energy efficient nodes to keep costs down.
GTX 970 at $329 was an outlier and the lowest x70 product of all time, on a trailing node (28nm again after 20nm fell through). GTX 670 launched at $399 for a similarly sized die over 10 years ago. GTX 1070 launched at $449 7 years ago with another similarly-sized (~300mm2) die. Turing, Ampere, and Maxwell were all abnormally cheap due and large due to the trailing-node but you paid for this with worse efficiency. There has never been a x70 product launched at $299 and you are welcome to check this!
First x70 product: https://en.wikipedia.org/wiki/List_of_Nvidia_graphics_proces...
And yes I think the consensus is that nodes like 8nm probably are "good enough" especially if they are much cheaper, especially at the low end where PHY size is becoming a problem. PHYs don't shrink, so you can't scale a product arbitrarily small - the logic may shrink by 70% but those PHYs are just as big as ever. So there is a de-facto "minimum die size" that is ever worth producing, because there is a fixed PHY area that you simply cannot eliminate. And in a world where wafer costs are going up, that area costs more and more every generation.
A 3060 Ti 16GB wouldn't even need clamshell (it has 8 PHYs, 8x2GB per module=16GB) and could probably have hit $299 or $329 launch cost, if NVIDIA had gone down that road. And it's Good Enough for 1080p, and avoids some weird compromises that shake out of the need to trim PHY area.
AMD already did exactly this with the 7600 - which is held back on 6nm (N7 family) rather than using N5P (N5 family) like the rest of the RDNA2 lineup. Why? Cost.