Live data from Hacker News

Nvidia announces next-gen RTX 5090 and RTX 5080 GPUs

theverge.com

721–730 of 776 posts

Re: Nvidia announces next-gen RTX 5090 and RTX 5080 GPUs

#721
post #716

Earlier quoted context omitted.

> AVX512 was that strategy’s zenith and it was a disaster for them since they had to cut clock speeds when AVX-512 operations ran while AMD was able to implement them without any apparent loss in clock speed. AMD up until zen 5 didn't have a full AVX-512 support so not exactly a fair comparison. Intel designs don't suffer from that issue AFAIU for couple of iterations already. But I agree with you, I always thought a…

AVX-512 is around a dozen different ISA extensions. AMD implemented the base AVX-512 and more with Zen 4. This was far more than Intel had implemented in skylake-X where their problems started. AMD added even more extensions with Zen 5, but they still do not have the full AVX-512 set of extensions implemented in a single CPU and neither does Intel. Intel never implemented every single AVX-512 extension in a single CP…

Ok, fair enough, I didn't explain myself very well. What I more specifically meant is that AMD up until zen5 could not

  (1) drive 2x AVX-512 computations
  (2) handle 2x AVX-512 memory loads + 1x AVX-512 memory store
in the same clock.

The latter makes a big impact wrt available memory BW per core, at least when it comes to the workloads whose data is readily available in L0 cache. Intel in these experiments is crushing AMD by a large factor simply because their memory controller design is able to sustain 2x64B loads + 1x64B stores in the same clock. E.g. 642 GB/s (Golden Cove) vs 334 GB/s (zen4) - this is a big difference and this is something that Intel had for ~10 years whereas AMD was able to solve this with zen5, basically only with the end of 2024.

Former one limits the theoretical FLOPS/core capabilities since single AVX-512 FMA operation in zen4 is implemented as two AVX2 uops occupying both FMA slots per clock. This is also big and, again, this is something where Intel had a lead up until zen5.

Wrt downclocking issues, they had a substantial impact with Skylake implementation but with Ice Lake this was a solved issue and this was in 2019. I'm cool with having ~97% of max freq budget available with heavy AVX-512 workloads.

OTOH AMD is also very thin with this sort of information and some experiments show that turbo boost clock frequency on zen4 lowers from one CCD to another CCD [1]. It seems like zen5 exhibits similar behavior [2].

So, although AMD is displaying continuous innovation for the past several years this is only because they had a lot to improve. Their pre-zen (2017) designs were basically crap and could not compete with Intel who OTOH had a very strong CPU design for decades.

I think that the biggest difference in CPU core design really is in the memory controller - this is something Intel will need to find an answer to since AMD matched all the Intel strengths that it was lacking with zen5.

[1] https://chipsandcheese.com/p/amds-zen-4-part-3-system-level-...

[2] https://chipsandcheese.com/p/amds-ryzen-9950x-zen-5-on-deskt...

Re: Nvidia announces next-gen RTX 5090 and RTX 5080 GPUs

#722
post #90
post #14

Even though they are all marketed as gaming cards, Nvidia is now very clearly differentiating between 5070/5070 Ti/5080 for mid-high end gaming and 5090 for consumer/entry-level AI. The gap between xx80 and xx90 is going to be too wide for regular gamers to cross this generation.

How will a 5090 compare against project digits? now that they're both in the front page :)

Don't forget that you can link for example two 'Digits' together (~256 GB) if you want to run even larger models or have larger context size. That is 2x$3000 vs 8x$2000.

This will make it possible for you to run models up to 405B parameters, like Llama 3.1 405B at 4bit quant or the Grok-1 314B at 6bit quant.

Who knows, maybe some better models will be released in the future which are better optimized and won't need that much RAM, but it is easier to buy a second 'Digits' in comparison to building a rack with 8xGPUs. For example, if you look at the latest Llama models, Meta states: 'Llama 3.3 70B approaches the performance of Llama 3.1 405B'.

To interfere with Llama3.3-70B-Instruct with ~8k context length (without offloading), you'd need: - Q4 (~44GB): 2x5090; 1x 'Digits' - Q6 (~58GB): 2x5090; 1x 'Digits' - Q8 (~74GB): 3x5090; 1x 'Digits' - FP16 (~144GB): 5x5090; 2x 'Digits'

Let's wait and see which bandwidth it will have.

Re: Nvidia announces next-gen RTX 5090 and RTX 5080 GPUs

#723
post #90

Earlier quoted context omitted.

How will a 5090 compare against project digits? now that they're both in the front page :)

Don't forget that you can link for example two 'Digits' together (~256 GB) if you want to run even larger models or have larger context size. That is 2x$3000 vs 8x$2000. This will make it possible for you to run models up to 405B parameters, like Llama 3.1 405B at 4bit quant or the Grok-1 314B at 6bit quant. Who knows, maybe some better models will be released in the future which are better optimized and won't need t…

> bandwidth

Speculation has it at ~5XXgb/s.

agreed on the memory.

if I can I'll get a few but I fear they'll sell out immediately

Re: Nvidia announces next-gen RTX 5090 and RTX 5080 GPUs

#724

Similar CUDA core counts for most SKUs compared to last gen (except in the 5090 vs. 4090 comparison). Similar clock speeds compared to the 40-series. The 5090 just has way more CUDA cores and uses proportionally more power compared to the 4090, when going by CUDA core comparisons and clock speed alone. All of the "massive gains" were comparing DLSS and other optimization strategies to standard hardware rendering. Som…

I started thinking today, when Nvidia seemingly keeps just magically increasing performance every two years, that they eventually have to "intel" themselves, where they haven't made any real architectural improvements in ~10 years and just suddenly power and thermals don't scale anymore and you have six generations of turds that all perform essentially the same, right?

As long as TSMC keeps improving die size it will keep getting incremental improvements. These power/thermal improvements are not really that much up to nvidia.

The intel problem was that their foundries couldn't improve the die size while the other foundries kept improving theirs. But technically nvidia can switch foundry if another one proves better than TSMC even though that doesn't seem likely (at least without a major breakthrough not capitalized by ASML).

Re: Nvidia announces next-gen RTX 5090 and RTX 5080 GPUs

#725
post #488

Earlier quoted context omitted.

If anyone thinks they are having laggier controls or losing latency off of single frames I have a bridge to sell them. A game running at 60 fps averages around ~16 ms and good human reaction times don’t go much below 200ms. Users who “notice” individual frames are usually noticing when a single frame is lagging for the length of several frames at the average rate. They aren’t noticing anything within the span of an a…

you’re conflating reaction times and latency perception. these are not the same. humans can tell the difference down to 10ms, perhaps lower. if you added 200ms latency to your mouse inputs, you’d throw your computer out the of the window pretty quickly.

yeah the "distance between frames" latency is just one overhead, everything adds up until you get real latency. 10ms for your wireless mouse then 3ms for your I/O hardware then 5ms for the game engine to process your input then 20ms for the graphics pipeline and so on and on.

30 FPS is 33.33333 MS 60 FPS is 16.66666 MS 90 FPS is 11.11111 MS 120 FPS is 8.333333 MS 140 FPS is 7.142857 MS 144 FPS is 6.944444 MS 180 FPS is 5.555555 MS 240 FPS is 4.166666 MS

Going from 30fps to 120fps is 25ms which is totally 100% noticeable even for layman (I actually tested this with my girlfriend, she could tell between 60fps and 120fps as well), but these generated frames from DLSS don't help with this latency _at all_.

Although the nVidia Reflex technology can help with this kind of latency in some situations in some non quantifiable ways.

Re: Nvidia announces next-gen RTX 5090 and RTX 5080 GPUs

#726

Earlier quoted context omitted.

4K alone is not enough to define spatial resolution. You also need take into account physical dimensions. DPI is a better way to describe spatial resolution. Anything better than 200 DPI is good, better than 300 is awesome. Unfortunately, there are no 4K displays with 200+ DPI on the market. If you want high DPI you either need pick glossy 5k@27" or go to 6k/8k.

of course "normal viewing distances" is always implied when talking about monitors. And if you REALLY want to get pedantic you need to talk about pixels per degree. The human eye can see about 60. according to the very handy site https://qasimk.io/screen-ppd/ a 27" 1080p screen has 37ppd at 2 feet. a 42" 4k screen has 51ppd at 2 feet. a 27" 8k screen has 147ppd at 2 feet which is just absurd. You have to get to 6 inc…

> The human eye can see about 60

I cannot brag with sharp eyesight, but I can definitely tell difference between 4k@27" at 60cm = 73PPD and 5k@27" at 60cm = 97PPD. Text is much crisper on the latter.

I've also compared Dell 8k to 6k. There is a still a difference, but it is not that big.

Re: Nvidia announces next-gen RTX 5090 and RTX 5080 GPUs

#727

Earlier quoted context omitted.

I was demonstrating the Apples to Oranges comparison. If they were both free no one would pick DLSS. It shows Rasterizing is preferable. So comparing Rasterizing performance to DLSS performance is dishonest.

Except that if rendering was magically free... why not just pathtrace everything? DLSS might not be as good as pure unlimited pathtracing, but for a given budget it might be better than rasterization alone.

I agree it’s worth the trade off. I use upscalers a lot.

I’m saying that it’s different enough that you shouldn’t compare the two.

Re: Nvidia announces next-gen RTX 5090 and RTX 5080 GPUs

#728

Earlier quoted context omitted.

I started thinking today, when Nvidia seemingly keeps just magically increasing performance every two years, that they eventually have to "intel" themselves, where they haven't made any real architectural improvements in ~10 years and just suddenly power and thermals don't scale anymore and you have six generations of turds that all perform essentially the same, right?

I mean it's like 1/6 of their revenue now and will probably keep sliding in importance over the datacenter. No real competition no matter how we would wish. AMD seems to have given up on the high end and Intel is focusing on the low end (for now, unless they cancel it in the next year or so).

From what I've seen they've targeted the low end in price, but solid mid-range in performance. It's hard to know if that's a strategy to get started (likely) with price increases down the road or they're really that competitive.

Intel's iGPUs were low end. Battlemage looks firmly mid-range at the moment with between 4060/4070 performance in a lot of cases.

Re: Nvidia announces next-gen RTX 5090 and RTX 5080 GPUs

#729

Earlier quoted context omitted.

of course "normal viewing distances" is always implied when talking about monitors. And if you REALLY want to get pedantic you need to talk about pixels per degree. The human eye can see about 60. according to the very handy site https://qasimk.io/screen-ppd/ a 27" 1080p screen has 37ppd at 2 feet. a 42" 4k screen has 51ppd at 2 feet. a 27" 8k screen has 147ppd at 2 feet which is just absurd. You have to get to 6 inc…

> The human eye can see about 60 I cannot brag with sharp eyesight, but I can definitely tell difference between 4k@27" at 60cm = 73PPD and 5k@27" at 60cm = 97PPD. Text is much crisper on the latter. I've also compared Dell 8k to 6k. There is a still a difference, but it is not that big.

"Much crisper"

You must have exceptional eyesight.

Re: Nvidia announces next-gen RTX 5090 and RTX 5080 GPUs

#730
post #716

Earlier quoted context omitted.

AVX-512 is around a dozen different ISA extensions. AMD implemented the base AVX-512 and more with Zen 4. This was far more than Intel had implemented in skylake-X where their problems started. AMD added even more extensions with Zen 5, but they still do not have the full AVX-512 set of extensions implemented in a single CPU and neither does Intel. Intel never implemented every single AVX-512 extension in a single CP…

Ok, fair enough, I didn't explain myself very well. What I more specifically meant is that AMD up until zen5 could not (1) drive 2x AVX-512 computations (2) handle 2x AVX-512 memory loads + 1x AVX-512 memory store in the same clock. The latter makes a big impact wrt available memory BW per core, at least when it comes to the workloads whose data is readily available in L0 cache. Intel in these experiments is crushing…

System memory is not able to sustain such memory bandwidth so it seems like a moot point to me. Intel’s CPUs reportedly cannot sustain such memory bandwidth even when it is available:

https://www.ixpug.org/images/docs/ISC23/McCalpin_SPR_BW_limi...

Post reply on HN