Live data from Hacker News

Nvidia is proposing a beast of a CPU system for Windows PCs

twitter.com

271–280 of 581 posts

Re: Nvidia is proposing a beast of a CPU system for Windows PCs

#271

The Unified Memory pool is what will continue to be the “game changer” in systems architecture, especially outside of data centers. The reality is even cutting edge games and consumer workloads don’t actually take full use of the PCIe bandwidth of the GPU or the bandwidth of its GDDR memory. Even local AI use cases don’t substantially or meaningfully benefit from faster memory, at least to average consumers. A unifie…

What is the difference between unified memory and shared memory?

Shared memory existed since the first CPU with an embedded GPU came to market and you could set in BIOS how much memory goes to what component.

I do have an opinion about how unified memory could be different, but I want a proper explanation.

Re: Nvidia is proposing a beast of a CPU system for Windows PCs

#272
post #249
post #213

Earlier quoted context omitted.

Unified memory is only a feature because NVidia so aggressively uses VRAM for market segmentation. The 5090 ($2k MSRP but realistically $3-3.5k) is almost the same as the RTX 6000 Pro (~$10k). Same memory bandwidth (1800GB/s). Slightly different CUDA cores (21k vs 24k). Big difference? VRAM (32GB vs 96GB). NVidia ultimately doesn't want to upset this segmentation so the RTX Spark will never undermine their other offe…

I have so many questions… Since Apple already sells unified memory systems, what is the market opportunity you envision? Do you see Nvidia and Apple as competitors, and how? (And I’m not suggesting they’re not, necessarily, but I want to hear where you’re coming from, and they do have very different markets.) Hasn’t Apple used storage size (RAM & disk) for market segmentation for decades? And how does a machine with…

I'm not the person you're replying to, but I wholeheartedly agree with them...

Quick background: doing AI inference requires three things. Lots of memory, lots of memory bandwidth, and of course plenty of compute that has access to that memory.

Quick reference: nVidia 5090 has 1,792 GB/sec bandwidth. 3090 gets about 1000 GB/sec. DGX Spark and AMD 395 whatever get about 275 GB/sec.

Apple M1 Max gets 400GB/sec, M5 Max gets 614GB/sec. Ultra variants get 2x that bandwidth, base variants get 1/2 that bandwidth. However... their compute is rather weak.

Right now, Apple's offerings are juuuuuust fast enough to run dense 27B models at usable speeds at like, 10% of the performance/watt of nVidia. They're world-leading general purpose CPUs but not killer GPUs.

By all accounts, these Windows PCs nVidia is touting seem to have DGX Spark like performance, which is less than impressive. Same with the upcoming AMD AI-oriented consumer stuff.

The other context here is that running your own AI at home is just starting to become feasible in terms of open model availability and the ability to run it at usable speeds. Many are interested in it for reasons of privacy, security, and cost certainty vs. buying tokens.

    Since Apple already sells unified memory systems, what 
    is the market opportunity you envision?
nVidia and AMD can't make their consumer offerings too good at AI, because that risks interfering with their higher-margin data center sales.

(And, let's face it. Even if nVidia did release a 6090 with 64-128GB of memory for an affordable price, consumers wouldn't get their hands on them anyway because people would just start filling data centers with them)

So.

Now you see Apple's opportunity, right? No data center sales to interfere with. No relationship with nVidia or AMD to worry about.

They could choose to make an absolute beast of a home AI machine. The M5 Ultra, if announced, might be that. It's admittedly a niche market, but people are already buying 64GB+ Macs faster than Apple can make them and they're fetching high prices on the used market as well.

The only real questions are if this market is even something Apple would find time to care about, and if they could secure enough DRAM to make a go at it. They are enormous obviously but they're feeling the RAM pinch just like everybody.

Re: Nvidia is proposing a beast of a CPU system for Windows PCs

#273
post #249
post #213

Earlier quoted context omitted.

Unified memory is only a feature because NVidia so aggressively uses VRAM for market segmentation. The 5090 ($2k MSRP but realistically $3-3.5k) is almost the same as the RTX 6000 Pro (~$10k). Same memory bandwidth (1800GB/s). Slightly different CUDA cores (21k vs 24k). Big difference? VRAM (32GB vs 96GB). NVidia ultimately doesn't want to upset this segmentation so the RTX Spark will never undermine their other offe…

I have so many questions… Since Apple already sells unified memory systems, what is the market opportunity you envision? Do you see Nvidia and Apple as competitors, and how? (And I’m not suggesting they’re not, necessarily, but I want to hear where you’re coming from, and they do have very different markets.) Hasn’t Apple used storage size (RAM & disk) for market segmentation for decades? And how does a machine with…

Apple offers relatively affordable options for a high-memory workstation that uses unified memory. They previously offered 256/512GB Mac Studios (both discontinued). Because of this they can keep larger models in memory.

BUT you just can't compete with NVidia performance for LLM workloads (mostly inference) for two reasons:

1. The memory bandwidth just can't compete with a 5090 (1800GB/s). The best current Mac is ~900GB/s. That directly caps tokens/sec and might be manageable but there's another problem; and

2. The raw FLOPS just can't compete with even a 5090. It probably needs to natively support FP4/FP8 to at least maintain a number format parity with NVidia. But beside that, NVidia just has more raw FLOPS.

According to Google, an M5 Max does ~70 FP16 TFLOPS while a 5090 does 380. If Apple can close that gap to at least be competitive and also hold larger models in shared VRAM, that would be a competitive advantage and it would directly attack NVidia's market segmentation.

The Mac Studio last came out March last year. So we may get an update in Q3. Many are pinning their hopes on this. But it might not happen until next year. When it was released the M4 was the state of the art and it came with either the M4 Max or M3 Ultra (which, as I understand it, is basically 2 M3s stuck together, kind of). What people are hoping for is an M5 Ultra with >1000GB/s of memory bandwidth, ideally 200+ FP16 TFLOPS and hopefully FP4/FP4 support.

You can chain Mac Studios together into a cluster with TB5 too.

But it's reasonably likely that the next Mac Studio will be only incrementally better than the last generation.

Re: Nvidia is proposing a beast of a CPU system for Windows PCs

#274
post #65

good to know, hope the price will be affordable, having a pc becoming a luxury :)

I’m not sure if you’re aware but there is a supply chain shortage for pretty much everything needed for a PC that isn’t expected to be solved this year or next year. There is no way that can be affordable

or moreso, the available supply has been eaten by rampant speculation, and hyperscalers have overpurchased vs the datacenters they can actually get built and power

Re: Nvidia is proposing a beast of a CPU system for Windows PCs

#275

Earlier quoted context omitted.

You must be unaware that System76 was already selling 192GB machines, mac studios used to be 512GB max. The only reason why we don’t have them anymore is that we are in RAM shortage.

Those 192GB aren't unified memory though. 128GB on Mac or 395 can be used by both CPU and GPU. It's the GPU + large memory that opens up fast local LLM inteference.

Yes, true. But if we had the ability to buy that much RAM in the laptop, everyone would be looking in that direction. Until this thing discussed here comes to the market, “we didn’t have computers with unified 128GB RAM either” (except of macs).

Re: Nvidia is proposing a beast of a CPU system for Windows PCs

#276

How much is this supposed to cost, fully populated with 128GB of RAM? How much would this laptop cost? It's not that the NVidia chip has that much RAM built in, after all. It's that it can address that much. RAM is sold separately.

No. It’s all integrated. Not something you can buy separately or upgrade later.

See [1]. There's not 128GB of on-chip memory. "Integrated" memory in this context means that the GPUs and CPUs all use the same memory. There are on-chip caches, of course.

[1] https://www.nvidia.com/en-us/products/rtx-spark/

Re: Nvidia is proposing a beast of a CPU system for Windows PCs

#277
post #212

The Qualcomm Snapdragon X2 Elite Extreme trounces Nvidia's chip in single core CPU performance. It beats Intel and AMD's best, too. It has unified memory. It's the only CPU in the same league as Apple's M-series in both CPU performance and power efficiency. And it's available in laptops today, not later this year. People are sleeping on Qualcomm.

Why do people care so much about single core performance? We are all professionals here and I bet most of our workloads are multi core. I get that these new arm chips from Apple and Qualcomm are great at one thing at a time, but for professional workloads high end x64 chips still cannot be beaten on the desktop.

Single core performance is the biggest factor for most day-to-day use of a computer, the stuff I do on a laptop. It's more important than peak multi core performance for web browsing and games. I only care about multi core performance when I'm compiling, and I usually do heavy compiles on a remote machine rather than on my laptop.

Re: Nvidia is proposing a beast of a CPU system for Windows PCs

#278
post #15

"I am not sure how many people will run AI models locally. It still seems like a niche application to me. However, it will make decent machines to play video games." I don't know who will be the winner but with some of the recent releases from gemma it seems more probable that you may run some models locally if only from a cost perspective, not even considering business security. Not sure how this type of architectur…

> not even considering business security.

I suspect personal privacy and need to run AI workflows to handle the litany of administration tasks of a household will be what result in regular need for local AI.

Apple is already out front with this on a personal, individual level, but they are not obviously headed toward multiuser/family-level ~biz admin with a persistent server running local LLM.

Re: Nvidia is proposing a beast of a CPU system for Windows PCs

#279

Earlier quoted context omitted.

That’s too strong of an assertion. Local models aren’t deterministically equivalent in capabilities to foundation models. Home computers are turing complete; just like a mainframe. They are just slower. Often not slower enough to matter.

Most people are ok with slower. An AI that lets you edit a family picture, in say 30 seconds, locally is preferable to one that is instantaneous but requires you to submit that picture to examination/storage/training/sale in someone else's AI ecosystem. If i want to crop my ex out of family photos, i should not have to first give that photo to Microsoft. If want an LLM to write a book report for me, i dont want it al…

It’s completely technically possible to have cloud services where customer data is opaque to the provider. Some of Apple’s services are like this already, for example.

I think there’s a sweet spot currently with munging your data blindly on the server so that your client device battery still lasts all day.

Meanwhile Apple and others push on with making client side models more efficient so that eventually the server costs and complexities go away.

Re: Nvidia is proposing a beast of a CPU system for Windows PCs

#280

Earlier quoted context omitted.

> This is the 2026 edition of Ken Olsen: "There is no reason anyone would want a computer in their home" Digging into this: > In conclusion, there is evidence that Ken Olsen did doubt the need for computers in the home, but the evidence is based primarily on the testimony of David Ahl who was perturbed when the personal computer project he championed at DEC was not supported by Olsen in 1974. > Olsen’s resistance may…

That's exactly the point. Until recently, AI models that could run on home machines were so bad that it was very hard to imagine anyone wanting to. And, like the overly large machines of 1977, models are getting faster, leaner, and better. It's happening a lot quicker, though.

This is why I'm bearish on Anthropic, OpenAI, and friends. I am not confident that we will continue to see the same pace of improvement in frontier model capabilities as we have seen over the past year or two - not using similar mathematics at least. But I think that getting results that are close enough to the same standard to be a realistic substitute but in a model small enough to run locally may well happen quite quickly. And if it does - where is the moat to defend these AI organisations with their astronomical budgets when they're already starting to price more realistically and that's already killing a lot of the hype they've enjoyed until very recently? They have an accidental moat because they bought up the global supply chain for storage but that surely isn't going to last once the data centres to hold that storage are becoming liabilities.
Post reply on HN