Earlier quoted context omitted.
Ok, fair enough, I didn't explain myself very well. What I more specifically meant is that AMD up until zen5 could not (1) drive 2x AVX-512 computations (2) handle 2x AVX-512 memory loads + 1x AVX-512 memory store in the same clock. The latter makes a big impact wrt available memory BW per core, at least when it comes to the workloads whose data is readily available in L0 cache. Intel in these experiments is crushing…
System memory is not able to sustain such memory bandwidth so it seems like a moot point to me. Intel’s CPUs reportedly cannot sustain such memory bandwidth even when it is available: https://www.ixpug.org/images/docs/ISC23/McCalpin_SPR_BW_limi...
Nvidia announces next-gen RTX 5090 and RTX 5080 GPUs
731–740 of 776 posts
Re: Nvidia announces next-gen RTX 5090 and RTX 5080 GPUs
#732Earlier quoted context omitted.
Not anymore. FSR4 is AMD only, and only the new RDNA4 GPUs.
I have seen AMD's PR materials for RDNA4, and as far as I can tell, they do not say anywhere anything like that. People read too much into "designed for RDNA4".
Why would they write that on their marketing slides?
Re: Nvidia announces next-gen RTX 5090 and RTX 5080 GPUs
#733Earlier quoted context omitted.
If you want to run LLMs buy their H100/GB100/etc grade cards. There should be no expectation that consumer grade gaming cards will be optimal for ML use.
> There should be no expectation that consumer grade gaming cards will be optimal for ML use. And yet it just so happens they work effectively the same. I've done research on an RTX 2070 with just 8 GB VRAM. That card consistently met or got close to the performance of a V100 albeit with less vram. Why indicate people shouldn't use consumer cards? It's dramatically (like 10x-50x) cheaper. Is machine learning only for…
Re: Nvidia announces next-gen RTX 5090 and RTX 5080 GPUs
#734Re: Nvidia announces next-gen RTX 5090 and RTX 5080 GPUs
#735Earlier quoted context omitted.
> There should be no expectation that consumer grade gaming cards will be optimal for ML use. And yet it just so happens they work effectively the same. I've done research on an RTX 2070 with just 8 GB VRAM. That card consistently met or got close to the performance of a V100 albeit with less vram. Why indicate people shouldn't use consumer cards? It's dramatically (like 10x-50x) cheaper. Is machine learning only for…
If you can do research on a mid tier consumer card then more power to you. I'm specifically referencing the people who are complaining that the specs on consumer video game GPUs are not good for ML work. Like theres just no reasonable expectation that they will be.
Which may be true although there are more differences than just VRAM and I assume those market segments have different perceptions of the real value Gamers want it cheaper/faster, institutions want it closer to state of the art, more robust to lengthy workloads (as in year long training sessions), and better support from nvidia. Among other things.
Re: Nvidia announces next-gen RTX 5090 and RTX 5080 GPUs
#736Earlier quoted context omitted.
PC power supplies already support 240V. Their connectors can take 120V or 240V.
Yes, but a standard household wall socket in the US supplies 120V @ 15A, for a max continuous power of 1.4 kW or so. So typical power supplies are only designed to draw up to that much power, or less. If someone made a PC power supply designed to plug into a NEMA 14-50 you could run a lot of GPUs! And generate a lot of heat!
https://www.amazon.com/IronBox-Electric-Connector-Power-Cord...
https://www.amazon.com/14-50P-6-15R-Adapter-Adaptor-Charger/...
As long as the PSU has proper overcurrent protection, you could get away with saying it is designed for this. I suspect you meant designed for higher power draw rather than merely designed to be able to be plugged into the receptacle, but your remark was ambiguous.
Usually, the way people do things to get higher power draw is that they have a power distribution unit that provides C14 receptacles and plugs into a high power outlet like this:
https://www.apc.com/us/en/product/APDU9981EU3/apc-rack-pdu-9...
Then they plug multiple power supplies into it. They are actually able to use the full available AC power this way.
A (small) problem with scaling PSUs to the 50A (40A continuous) that NEMA 14-50 provides is that there is no standard IEC connector for it as far as I know. The common C13/C14 connectors are limited to 10A. The highest is C19/C20 for 16A, which is used by the following:
https://seasonic.com/atx3-prime-px-2200/
https://seasonic.com/prime-tx/
If I read the specification sheets correctly, the first one is exclusively for 200-240VAC while the second one will go to 1600W off 120V, which is permitted by NEMA 5-15 as long as it is not a continuous load.
There is not much demand for higher rated PSUs in the ATX form factor most here would want, but companies without brand names appear to make ones that go up to 3.6kW:
https://www.amazon.com/Supply-Bitcoin-Miners-Mining-180-240V...
As for even higher power ratings, there are companies that make them in non-standard form factors if you must have them. Here is one example:
https://www.infineon.com/cms/en/product/promopages/AI-PSU/#1...
Re: Nvidia announces next-gen RTX 5090 and RTX 5080 GPUs
#737Earlier quoted context omitted.
But MPEG is lossy compression which means they are kind of a just a guess. That is why MPEG uses motion vectors. "MPEG uses motion vectors to efficiently compress video data by identifying and describing the movement of objects between frames, allowing the encoder to predict pixel values in the current frame based on information from previous frames, significantly reducing the amount of data needed to represent the v…
There's a real difference between a lossy approximation as done by video compression, and the "just a guess" done by DLSS frame generation. Video encoders have the real frame to use as a target; when trying to minimize the artifacts introduced by compressing with reference to other frames and using motion vectors, the encoder is capable of assessing its own accuracy. DLSS fundamentally has less information when gener…
Re: Nvidia announces next-gen RTX 5090 and RTX 5080 GPUs
#738Similar CUDA core counts for most SKUs compared to last gen (except in the 5090 vs. 4090 comparison). Similar clock speeds compared to the 40-series. The 5090 just has way more CUDA cores and uses proportionally more power compared to the 4090, when going by CUDA core comparisons and clock speed alone. All of the "massive gains" were comparing DLSS and other optimization strategies to standard hardware rendering. Som…
Re: Nvidia announces next-gen RTX 5090 and RTX 5080 GPUs
#739I'm really disappointed in all the advancement in frame generation. Game devs will end up relying on it for any decent performance in lieu of actually optimizing anything, which means games will look great and play terribly. It will be 300 fake fps and 30 real fps. Throw latency out the window.
It doesn't matter if that's through software or hardware improvements.
Re: Nvidia announces next-gen RTX 5090 and RTX 5080 GPUs
#740Earlier quoted context omitted.
System memory is not able to sustain such memory bandwidth so it seems like a moot point to me. Intel’s CPUs reportedly cannot sustain such memory bandwidth even when it is available: https://www.ixpug.org/images/docs/ISC23/McCalpin_SPR_BW_limi...
Not sure I understood you. You think that AVX-512 workload and store-load BW are irrelevant because main system memory (RAM) cannot keep up with the speed of CPU caches?
https://www.ixpug.org/images/docs/ISC23/McCalpin_SPR_BW_limi...
Your 642 GB/s figure should be for a single Golden Cove core, and it should only take 3 Golden Cove cores to saturate the 1.6 TB/sec HBM2e in Xeon Max, yet internal bottlenecks prevented 56 Golden Cove cores from reaching the 642 GB/s read bandwidth you predicted a single core could reach when measured. Peak read bandwidth was 590 GB/sec when all 56 cores were reading.
According to the slides, peak read bandwidth for a single Golden Cove core in the sapphire rapids CPU that they tested is theoretically 23.6GB/sec and was measured at 22GB/sec.
Chips and Cheese did read bandwidth measurements on a non-HBM2e version of sapphire rapids:
https://chipsandcheese.com/p/a-peek-at-sapphire-rapids
They do not give an exact figure for multithreaded L3 cache bandwidth, but looking at their chart, it is around what TACC measured for HBM2e. For single threaded reads, it is about 32 GB/sec from L3 cache, which is not much better than it was for reads from HBM2e and is presumably the effect of lower latencies for L3 cache. The Chips and Cheese chart also shows that Sapphire Rapids reaches around 450 GB/sec single threaded read bandwidth for L1 cache. That is also significantly below your 642 GB/sec prediction.
The 450 GB/sec bandwidth out of L1 cache is likely a side effect of the low latency L1 accesses, which is the real purpose of L1 cache. Reaching that level of bandwidth out of L1 cache is not likely to be very useful, since bandwidth limited operations will operate on far bigger amounts of memory than fit in cache, especially L1 cache. When L1 cache bandwidth does count, the speed boost will last a maximum of about 180ns, which is negligible.
What bandwidth CPU cores should be able to get based on loads/stores per clock and what bandwidth they actually get are rarely ever in agreement. The difference is often called the Von Neumann bottleneck.