I’m getting more and more convinced that we will end up running LLMs in our personal computers. Which makes me wonder where Anthropic/OpenAIs moats will come from.
Convince me 1. in order to run LLMs, especially the best ones, you need complicated devices which are expensive 2. if you buy one for your personal use, you are probably not going to utilize it all the time and it will be idle a lot It seems to me that it will always be more economical that the LLM-running devices are in a datacenter where it is easier to make sure they are always utilized
Nvidia RTX Spark
101–110 of 437 posts
Re: Nvidia RTX Spark
#102I’m getting more and more convinced that we will end up running LLMs in our personal computers. Which makes me wonder where Anthropic/OpenAIs moats will come from.
Re: Nvidia RTX Spark
#103I’m getting more and more convinced that we will end up running LLMs in our personal computers. Which makes me wonder where Anthropic/OpenAIs moats will come from.
Convince me 1. in order to run LLMs, especially the best ones, you need complicated devices which are expensive 2. if you buy one for your personal use, you are probably not going to utilize it all the time and it will be idle a lot It seems to me that it will always be more economical that the LLM-running devices are in a datacenter where it is easier to make sure they are always utilized
Local models today are fine for a lot of mundane tasks and will continue to be so. The use cases where paying for frontier models is worth it, will continue to shrink for folks not doing frontier work.
Re: Nvidia RTX Spark
#104I’m getting more and more convinced that we will end up running LLMs in our personal computers. Which makes me wonder where Anthropic/OpenAIs moats will come from.
Convince me 1. in order to run LLMs, especially the best ones, you need complicated devices which are expensive 2. if you buy one for your personal use, you are probably not going to utilize it all the time and it will be idle a lot It seems to me that it will always be more economical that the LLM-running devices are in a datacenter where it is easier to make sure they are always utilized
Think of it like having a graphics card at home versus using a cloud gaming stream? Technically subscribing to GeForce is much cheaper up front than getting a card, but people still do that. So will the audience of people running agents at home be as large as PC gaming? I think that's kind of plausible.
Re: Nvidia RTX Spark
#105I’m getting more and more convinced that we will end up running LLMs in our personal computers. Which makes me wonder where Anthropic/OpenAIs moats will come from.
Re: Nvidia RTX Spark
#106So they have basically reused the same hardware as in the DGX Spark (GB10)... That chip isn't great for LLM inference actually. https://www.techpowerup.com/gpu-specs/gb10.c4342 https://www.nvidia.com/en-us/products/rtx-spark/
>That chip isn't great for LLM inference actually. Why do I have the feeling it's been intentionally made to be bad in order to get you on to their most pensive datacenter gear.
At this point, your cost-efficient options include used 3090s, "frankenrigs" using recycled data center cards, and a handful of "workstation" class cards, where the originally high margins and the long enterprise purchasing cycles have kept prices from going up too fast.
In contrast, a lot of these "personal" AI systems are basically a GPU-like core wired to larger amounts of slow RAM. Which is still semi-affordable. Generally speaking, they make for OK chatbots but extremely slow coding agents. Whereas you can run a modestly useful coding agent at reasonable speed on a 3090.
So yeah, a lot of these systems are bit scammy. But not because it's a secret conspiracy to protect data center cards. Rather, there simply isn't enough fast RAM in the entire world. So they'll flog you disappointly slow RAM instead.
TL;dr: Might be useful for some use cases, but benchmark very carefully.
Re: Nvidia RTX Spark
#107I’m getting more and more convinced that we will end up running LLMs in our personal computers. Which makes me wonder where Anthropic/OpenAIs moats will come from.
Convince me 1. in order to run LLMs, especially the best ones, you need complicated devices which are expensive 2. if you buy one for your personal use, you are probably not going to utilize it all the time and it will be idle a lot It seems to me that it will always be more economical that the LLM-running devices are in a datacenter where it is easier to make sure they are always utilized
2. Eventually we'll get to where local models that don't have sycophancy and slot-machine mechanics trained into them will perform better.
Re: Nvidia RTX Spark
#108So they have basically reused the same hardware as in the DGX Spark (GB10)... That chip isn't great for LLM inference actually. https://www.techpowerup.com/gpu-specs/gb10.c4342 https://www.nvidia.com/en-us/products/rtx-spark/
Basically the same tradeoff as macmini with unified memory.
Re: Nvidia RTX Spark
#109Earlier quoted context omitted.
The ConnectX 7 2x200 Gbps networking card in the DGX Spark alone is worth $700
To be fair the connectx-7 in the spark can't even push 2x200 Gbps since it is connected via 4 pcie lanes.
Re: Nvidia RTX Spark
#110First: > "Our goal is to deliver unmetered intelligence to every home and every desk with Windows," said Satya Nadella, chairman and head of Microsoft. Then: > However, Ian Fogg, Research Director at industry analyst firm FDM CCS Insight said the change was "likely to come with a significant price tag" and Nvidia would be targeting "those looking for workstation-class performance". So... not every desk with Windows.