Live data from Hacker News

Nvidia RTX Spark

nvidia.com

101–110 of 437 posts

Re: Nvidia RTX Spark

#101
post #98
post #92

I’m getting more and more convinced that we will end up running LLMs in our personal computers. Which makes me wonder where Anthropic/OpenAIs moats will come from.

Convince me 1. in order to run LLMs, especially the best ones, you need complicated devices which are expensive 2. if you buy one for your personal use, you are probably not going to utilize it all the time and it will be idle a lot It seems to me that it will always be more economical that the LLM-running devices are in a datacenter where it is easier to make sure they are always utilized

Just like cloud vs private server. It'll be based on use case.

Re: Nvidia RTX Spark

#102
post #92

I’m getting more and more convinced that we will end up running LLMs in our personal computers. Which makes me wonder where Anthropic/OpenAIs moats will come from.

First we need to actually still be employed, and have them at affordable price.

Re: Nvidia RTX Spark

#103
post #98
post #92

I’m getting more and more convinced that we will end up running LLMs in our personal computers. Which makes me wonder where Anthropic/OpenAIs moats will come from.

Convince me 1. in order to run LLMs, especially the best ones, you need complicated devices which are expensive 2. if you buy one for your personal use, you are probably not going to utilize it all the time and it will be idle a lot It seems to me that it will always be more economical that the LLM-running devices are in a datacenter where it is easier to make sure they are always utilized

It's inevitable. What might be a prosumer device today priced at 4000$ will be a regular consumer device in 10 years and models only get better.

Local models today are fine for a lot of mundane tasks and will continue to be so. The use cases where paying for frontier models is worth it, will continue to shrink for folks not doing frontier work.

Re: Nvidia RTX Spark

#104
post #98
post #92

I’m getting more and more convinced that we will end up running LLMs in our personal computers. Which makes me wonder where Anthropic/OpenAIs moats will come from.

Convince me 1. in order to run LLMs, especially the best ones, you need complicated devices which are expensive 2. if you buy one for your personal use, you are probably not going to utilize it all the time and it will be idle a lot It seems to me that it will always be more economical that the LLM-running devices are in a datacenter where it is easier to make sure they are always utilized

If there end up being useful workflows where you keep stuff running in the background or overnight that's one advantage, compared to a data center that might cut off your access during peak hours or etc.

Think of it like having a graphics card at home versus using a cloud gaming stream? Technically subscribing to GeForce is much cheaper up front than getting a card, but people still do that. So will the audience of people running agents at home be as large as PC gaming? I think that's kind of plausible.

Re: Nvidia RTX Spark

#105
post #92

I’m getting more and more convinced that we will end up running LLMs in our personal computers. Which makes me wonder where Anthropic/OpenAIs moats will come from.

It'll be just another round of the client-side vs server-side processing rounds. We've been through them, we will keep going through them.

Re: Nvidia RTX Spark

#106

So they have basically reused the same hardware as in the DGX Spark (GB10)... That chip isn't great for LLM inference actually. https://www.techpowerup.com/gpu-specs/gb10.c4342 https://www.nvidia.com/en-us/products/rtx-spark/

>That chip isn't great for LLM inference actually. Why do I have the feeling it's been intentionally made to be bad in order to get you on to their most pensive datacenter gear.

It's probably more that LLM inference speed comes from having a large amount of fast RAM. And fast RAM is brutally expensive right now.

At this point, your cost-efficient options include used 3090s, "frankenrigs" using recycled data center cards, and a handful of "workstation" class cards, where the originally high margins and the long enterprise purchasing cycles have kept prices from going up too fast.

In contrast, a lot of these "personal" AI systems are basically a GPU-like core wired to larger amounts of slow RAM. Which is still semi-affordable. Generally speaking, they make for OK chatbots but extremely slow coding agents. Whereas you can run a modestly useful coding agent at reasonable speed on a 3090.

So yeah, a lot of these systems are bit scammy. But not because it's a secret conspiracy to protect data center cards. Rather, there simply isn't enough fast RAM in the entire world. So they'll flog you disappointly slow RAM instead.

TL;dr: Might be useful for some use cases, but benchmark very carefully.

Re: Nvidia RTX Spark

#107
post #98
post #92

I’m getting more and more convinced that we will end up running LLMs in our personal computers. Which makes me wonder where Anthropic/OpenAIs moats will come from.

Convince me 1. in order to run LLMs, especially the best ones, you need complicated devices which are expensive 2. if you buy one for your personal use, you are probably not going to utilize it all the time and it will be idle a lot It seems to me that it will always be more economical that the LLM-running devices are in a datacenter where it is easier to make sure they are always utilized

Uploading your IP to the biggest IP thieves in human history seems bad idk.

2. Eventually we'll get to where local models that don't have sycophancy and slot-machine mechanics trained into them will perform better.

Re: Nvidia RTX Spark

#108

So they have basically reused the same hardware as in the DGX Spark (GB10)... That chip isn't great for LLM inference actually. https://www.techpowerup.com/gpu-specs/gb10.c4342 https://www.nvidia.com/en-us/products/rtx-spark/

It is great for inference for single user/single session. it is not replacement for graphical accelerator, that run several concurrent inference sessions in parallel.

Basically the same tradeoff as macmini with unified memory.

Re: Nvidia RTX Spark

#109
post #64

Earlier quoted context omitted.

The ConnectX 7 2x200 Gbps networking card in the DGX Spark alone is worth $700

To be fair the connectx-7 in the spark can't even push 2x200 Gbps since it is connected via 4 pcie lanes.

Technically it's connected via 8 PCIe gen 5 lanes (two 4x connections), allowing ~100Gbps per port.

Re: Nvidia RTX Spark

#110

First: > "Our goal is to deliver unmetered intelligence to every home and every desk with Windows," said Satya Nadella, chairman and head of Microsoft. Then: > However, Ian Fogg, Research Director at industry analyst firm FDM CCS Insight said the change was "likely to come with a significant price tag" and Nvidia would be targeting "those looking for workstation-class performance". So... not every desk with Windows.

The constant deliberate conflagration of LLMs with general intelligence is so grating.
Post reply on HN