Live data from Hacker News

Nvidia is proposing a beast of a CPU system for Windows PCs

twitter.com

551–560 of 581 posts

Re: Nvidia is proposing a beast of a CPU system for Windows PCs

#551
post #213

The Unified Memory pool is what will continue to be the “game changer” in systems architecture, especially outside of data centers. The reality is even cutting edge games and consumer workloads don’t actually take full use of the PCIe bandwidth of the GPU or the bandwidth of its GDDR memory. Even local AI use cases don’t substantially or meaningfully benefit from faster memory, at least to average consumers. A unifie…

Unified memory is only a feature because NVidia so aggressively uses VRAM for market segmentation. The 5090 ($2k MSRP but realistically $3-3.5k) is almost the same as the RTX 6000 Pro (~$10k). Same memory bandwidth (1800GB/s). Slightly different CUDA cores (21k vs 24k). Big difference? VRAM (32GB vs 96GB). NVidia ultimately doesn't want to upset this segmentation so the RTX Spark will never undermine their other offe…

It's also ECC ram but to be fair - yes quite overpriced. The RTX Pro line are basically what the Titan line used to be but way way more expensive.

Re: Nvidia is proposing a beast of a CPU system for Windows PCs

#553

Earlier quoted context omitted.

Memory is just one part. AMD has had offerings competitive to NVIDIA for quite some time, but nobody uses AMD cards. The biggest advantage with NVIDIA is CUDA.

> but nobody uses AMD cards AMD is selling every MI card it makes, and the market wants more of them.

They are only selling because Nvidia is hard to get, and something is better than nothing.

Re: Nvidia is proposing a beast of a CPU system for Windows PCs

#554
post #427

Earlier quoted context omitted.

The M1 isnt particularly good at inference, so pretty much every major current competitor with a 256+ bit unified memory system is better: AMD Strix Halo, NVIDIA DGX Spark, possibly Intel Panther Lake

Sure, but none of these shipped before the M1. That was the first chip I encountered that managed to do something useful without a discrete GPU.

Even in Apple land M1 isn't the first with unified memory - pretty much all intel on-chip GPU (Sandy Bridge and newer) count - it was even a reason for driver issues early in intel's new dedicated GPU lineup, as the drivers expected unified memory - but M1 is essentially modification of an iPad chip, and you can see "unified memory" there going all the way back to first chip Apple bought from Samsung to power iPhone 2G

Re: Nvidia is proposing a beast of a CPU system for Windows PCs

#555

Earlier quoted context omitted.

We might even decide to put 32GB of high-latency cache on the system board and then 12GB of throughput-optimized main memory close to the GPU. ;)

You meant a 128GB (instead of 12GB)? And yes, a L4 cache can be one way out of that problem. Another way is making the L3 cache lines wider and working the hell out of improving your management algorithm. It's not a theoretically impossible problem. It's also not something you can solve automatically with a bit more money or some simple decisions. It's possible this is the best architecture available, but it's not ce…

I mean 12GB, an amount that is typical in such a system today, which you can buy at any computer store.

Re: Nvidia is proposing a beast of a CPU system for Windows PCs

#556
post #464

Earlier quoted context omitted.

Not true. This is aimed squarely at the Strix Halo and Mac markets. It's basically just strictly better than the Strix, and it's not clear cut vs that Macs in any sort of blanket statement. My M5 Max 128gb MBP decodes faster than one of my Sparks, but the Spark's prefill is so much faster it can often answer the same query before the mac's prefill is finished. If you have large prompts, low cacheability, etc., a spar…

Fair, but I don’t see what case you have w this. Mind sharing? Seems niche to be both uncacheable and long context?

Anything where you're dealing with a large volume of records/documents. Lots of people are using these for large-scale digitization of documents - scanned stuff being OCR'ed and summarized, generating embeddings, etc. Large scale translation.

Anywhere where you might have a large backlog of data to work with can end up in this sort of situation.

Re: Nvidia is proposing a beast of a CPU system for Windows PCs

#557
post #501
post #411

Earlier quoted context omitted.

I'd bet on the inverse: China scaling DRAM production until the price crashes, and the whole US stock market that is propped on top of that scarcity going down with it.

Tarrifs.

Plenty of demand outside the US. Why would the hyperscalers not buy the chinese RAM for all of their datacenters across the world besides the US one?

Rising supply from China will impact prices even in countries where there are tariffs.

Re: Nvidia is proposing a beast of a CPU system for Windows PCs

#558
post #501

Earlier quoted context omitted.

Tarrifs.

Plenty of demand outside the US. Why would the hyperscalers not buy the chinese RAM for all of their datacenters across the world besides the US one? Rising supply from China will impact prices even in countries where there are tariffs.

The best Chinese RAM on the market is 50% larger and requires more power and thus emits more heat, as it is a 16nm feature size. If they can get to competitive sizes, then of course data centers will purchase it.

Re: Nvidia is proposing a beast of a CPU system for Windows PCs

#559

Earlier quoted context omitted.

I think much of the difficulty is just that, for example, the 1.8 TB/s of an RTX 5090 is a lot of bandwidth for a game to use. That's over 50,000 4k textures per second at 32bpp.

I agree with you in theory. A couple of points - that’s currently the most experience and high performing card on the market. Most people on steam are using an RTX 3060 which has more like 360GB/s. That’s a factor of 6. How do you design resource usage that scales with that amount of extremity? (We try to, fwiw). That spec is also a throughput measured per second whereas our frame rates are much higher than 1/s. At 6…

All good points.

Absolutely an RTX 3060 is a more normal gamer GPU than the 5090, but you're also not playing in 4k without DLSS on a 3060. Drop to the most common resolution on Steam (1080p), and turn on DLSS and you've basically cancelled out that 6x factor in bandwidth. Even if the 3060 had more bandwidth, it doesn't have enough processing power for native 4k gaming in typical games. So 360 GB/s is still a lot of bandwidth for the resolution most 3060 gamers are using.

Re: Nvidia is proposing a beast of a CPU system for Windows PCs

#560

Earlier quoted context omitted.

> LPCAMM and similar solutions exist, but have never been demonstrated running at speeds that match what the leading soldered memory systems are using; there's always been some speed penalty. LPCAMM2 supports up to 9600MT/s, which appears to be the same speed Apple is using. > I'm not sure we've ever seen a system demonstrated using LPCAMM or similar for a 512-bit bus Servers commonly use a 768-bit DDR5 memory bus pe…

> LPCAMM2 supports up to 9600MT/s, which appears to be the same speed Apple is using. The difference here is in what the standard defines on paper vs what is actually shipping in products and readily available off the shelf. Who's selling a whole system with LPCAMM2 certified for 9600MT/s? Intel's current-gen Panther Lake top of the line laptop chips are rated for 9600MT/s when using soldered LPDDR5x but only 7467MT/…

> Who's selling a whole system with LPCAMM2 certified for 9600MT/s?

The 9600MT/s modules are new and will probably be found at some point this year. Framework already sells LPCAMM2 at 8533MT/s with full validation:

https://knowledgebase.frame.work/what-drammemory-is-supporte...

> That puts the current Intel-with-LPCAMM2 supported memory speed at 1.5 years and counting lag behind Apple's shipping memory speeds.

It turns out Apple isn't getting 9600MT/s either. I assumed that soldering would be getting them at least what LPCAMM2 is rated for, but if you actually do the math, they're getting ~8500MT/s for their most expensive systems and ~7500MT/s for the others.

> Servers aren't anywhere close to 9600MT/s yet; Intel and AMD are at 6400MT/s.

Servers use conservative timings. EXPO memory kits above 6400MT/s are available for Threadripper with 8 channels. And again, these are using traditional DIMMs with longer traces rather than CAMM, but they're still managing an extremely wide bus with close to the same performance.

> The trace length advantages offered by LPCAMM2 don't necessarily mean the traces for the sixth or eighth channel would be short enough for 9600MT/s

CAMM modules use a compression fitting to attach the chips to the system board using approximately the same amount of space as the solder pads would for soldered chips. If you get to the point of having so many channels that the chips are in the way of the other chips then the soldered ones have the same problem.

> (which again, is not yet available even in a 128-bit configuration in shipping hardware).

A single LPCAMM2 module is a 128-bit bus. Every system that uses it has at least that.

> Maybe you could get to 512-bit with modules on the front and back of the board while maintaining trace lengths short enough to reach meaningfully higher speeds than regular DDR5, but so far nobody is doing that or even talking about it.

Nobody is really using a bus that wide with soldered memory either though, outside of the couple of Macs that start at ~$3500 and are getting the same speed Framework does with LPCAMM2.

Post reply on HN