I can't really see wide adoption of local LLMs unless prices really start to climb. It makes sense to use cheaper hosted smaller models like Sonnet or even Kimi but these won't run a Kimi-class model and that is really the floor for non-toy agentic tasks. Spending 5k to avoid a $20 subscription really only makes sense for niche security reasons.
I'd bet on the inverse: China scaling DRAM production until the price crashes, and the whole US stock market that is propped on top of that scarcity going down with it.
Nvidia is proposing a beast of a CPU system for Windows PCs
501–510 of 581 posts
Re: Nvidia is proposing a beast of a CPU system for Windows PCs
#502Re: Nvidia is proposing a beast of a CPU system for Windows PCs
#503Sorry for lazy question but some people here know this off top of their head so asking. Memory bandwidth of this chip? Last time I check an NVidia situation was for DGX Spark (the GB10 chip), it has regular LPDDR5X which by JEDEC standard cannot go beyond ~270 GB/sec, ie 8533 Mbit/s on a 256 lanes bus. So yeah Lemire seems to go "OMG unified memory, they're following Apple path..." ok, but Apple pulled off a much fas…
Re: Nvidia is proposing a beast of a CPU system for Windows PCs
#504The Unified Memory pool is what will continue to be the “game changer” in systems architecture, especially outside of data centers. The reality is even cutting edge games and consumer workloads don’t actually take full use of the PCIe bandwidth of the GPU or the bandwidth of its GDDR memory. Even local AI use cases don’t substantially or meaningfully benefit from faster memory, at least to average consumers. A unifie…
> Lets systems optimize utilization based on need, rather than be confined to specific pools The trouble with this is that the different types of memory have different characteristics. Latency for ordinary system memory is actually better than it is for GDDR, because GDDR is optimized for bandwidth. RTX 5090 has 1.8TB/s of memory bandwidth with a 512-bit memory bus. The same bus width for DDR5-9600 would have better…
Re: Nvidia is proposing a beast of a CPU system for Windows PCs
#505Sorry for lazy question but some people here know this off top of their head so asking. Memory bandwidth of this chip? Last time I check an NVidia situation was for DGX Spark (the GB10 chip), it has regular LPDDR5X which by JEDEC standard cannot go beyond ~270 GB/sec, ie 8533 Mbit/s on a 256 lanes bus. So yeah Lemire seems to go "OMG unified memory, they're following Apple path..." ok, but Apple pulled off a much fas…
This is the same chip and same memory. Only difference is it is going in a laptop, so will be more thermally limited.
Re: Nvidia is proposing a beast of a CPU system for Windows PCs
#506Earlier quoted context omitted.
You seriously think running LLM is the same thing as general computing?
It’s better, it’s useful even for those who don’t have a deep knowledge of computers. I’d expect more AI users than programmers, than ms-word users, than excel users.
Re: Nvidia is proposing a beast of a CPU system for Windows PCs
#507> I am not sure how many people will run AI models locally. It still seems like a niche application to me. Bill Gates had a quote some years ago... People have still not learned how fast we improve our tech and how much cheaper thing gets I guess :)
We had a thing called globalism that drastically reduced costs. Globalism right now is on life support. Given geopolitics, I don’t see how it’s going to survive.
Re: Nvidia is proposing a beast of a CPU system for Windows PCs
#508Are their enterprise orders slowing down? Why use precious maxed out fab capacity on consumer stuff when it could be an enterprise chip?
Re: Nvidia is proposing a beast of a CPU system for Windows PCs
#509cant wait til someone figures how to run Linux on one of these
Re: Nvidia is proposing a beast of a CPU system for Windows PCs
#510Earlier quoted context omitted.
Most people are ok with slower. An AI that lets you edit a family picture, in say 30 seconds, locally is preferable to one that is instantaneous but requires you to submit that picture to examination/storage/training/sale in someone else's AI ecosystem. If i want to crop my ex out of family photos, i should not have to first give that photo to Microsoft. If want an LLM to write a book report for me, i dont want it al…
It’s completely technically possible to have cloud services where customer data is opaque to the provider. Some of Apple’s services are like this already, for example. I think there’s a sweet spot currently with munging your data blindly on the server so that your client device battery still lasts all day. Meanwhile Apple and others push on with making client side models more efficient so that eventually the server c…
If asked to choose between photo editing done within 3s using cloud provider vs an average of 30s using local compute, most consumers will choose the former without hesitation.
Most users' usage is also going to fall nicely in the free tier of a typical freemium pricing model, like ChatGPT today.
People who talk endlessly about local inference have no idea about user workflows and usability.