Live data from Hacker News

Jamesob's guide to running SOTA LLMs locally

github.com

131–140 of 193 posts

Re: Jamesob's guide to running SOTA LLMs locally

#131

Earlier quoted context omitted.

> in 18-24 months when they cost significantly less ... going to need you to sit down for this one...

Say more. My expectation is that the current gen of gpus will start being replaced by the next gen, and then it may be possible to get used ones that are still within their useful life at lower prices. My expectation is also that memory vendors are likely to increase production, which will drive those prices down eventually. Maybe not over the next 18-24 months though.

Given the prices and shortages, I'd think people would keep and use the current gen stuff till it drops dead. It may not be as good but it's paid for and given the prices for next gen stuff, it's probably worth using for another cycle or two.

Re: Jamesob's guide to running SOTA LLMs locally

#133
post #49

Earlier quoted context omitted.

Stop trying to run them locally, folks. You don't own your fiber connection. So why try to own another rapidly depreciating, expensive, and annoying asset? Rent cloud GPUs! You get to participate in the ownership, data control, price control, and hacking culture without having to Frankenstein some hobbyist box that costs a ton, is distilled down to functional uselessness, and is a PITA to maintain.

> You don't own your fiber connection. So why try to own another rapidly depreciating, expensive, and annoying asset? Like a car? Because I don’t want to depend on uber or a taxi service. What does fiber have to do with anything? I don’t need the internet to run my local models.

Yes, "Like a car". LOL. You realise that many people in Europe and Asia do not own a car at all? Public transport, eBike / scooter, Tuk-Tuk, walk.

The local LLM "privacy" war had been already lost.

Re: Jamesob's guide to running SOTA LLMs locally

#134

Earlier quoted context omitted.

> in 18-24 months when they cost significantly less ... going to need you to sit down for this one...

Say more. My expectation is that the current gen of gpus will start being replaced by the next gen, and then it may be possible to get used ones that are still within their useful life at lower prices. My expectation is also that memory vendors are likely to increase production, which will drive those prices down eventually. Maybe not over the next 18-24 months though.

The only thing that diminishes the value of a GPU right now is unsupported features with outsized value during inference and/or training (like FP4 support) and it takes time for those features to actually take off

And labs are fully leaning into pricing for intelligence, so their margins are improving very quickly (which allows them to pay even more for existing compute)

I'd be shocked if current prices aren't the bottom for the next 18-24 months.

Re: Jamesob's guide to running SOTA LLMs locally

#135
post #63

What harness is the best for local LLMs? I've been researching optimizing local LLM agent harness performance with context/ tools. Quite the endeavor and would love to learn what users prefer for this type of workflow.

What's the technical reason we call call these a harness? Seems right but want to understand better.

Re: Jamesob's guide to running SOTA LLMs locally

#136

I play with local LLMs a lot. I've spent more on hardware than I should. I'm friends with a local group of people who have spent a lot more than I have. The warning I would have for everyone is to temper your expectations and read the fine print carefully. The big build in article starts off with a $40K budget and then includes 4 GPUs that are $12K each. For those doing the math, this build is going to cost more like…

All very true. Right now, running GLM 5.2 at its full BF16 quantization level needs 1.5 TB of VRAM. You can't run this locally at a usable speed for less than $250K or so, and frankly I'd be surprised if it could be done for less than $500K. The best NV4FP quant for 5.2 appears to be lukealonso's at https://huggingface.co/lukealonso/GLM-5.2-NVFP4 , and it is capable of good throughput (75-100 tps) without losing much…

Nah we’re in like desktop PCs in the 90s type days - bit clunky and maybe occasionally having to work out what an IRQ number is, but a long long way from hand-toggling switches just to get a “hello world” punched out onto paper tape. You can go to an Apple Store today and an hour later have your AI agent talking to LM Studio, and it configuring your MacBook to code and do useful work while also running a diffusion model in the background. Slowly, but not “hand toggling hex switches” slow.

Re: Jamesob's guide to running SOTA LLMs locally

#137

Earlier quoted context omitted.

All very true. Right now, running GLM 5.2 at its full BF16 quantization level needs 1.5 TB of VRAM. You can't run this locally at a usable speed for less than $250K or so, and frankly I'd be surprised if it could be done for less than $500K. The best NV4FP quant for 5.2 appears to be lukealonso's at https://huggingface.co/lukealonso/GLM-5.2-NVFP4 , and it is capable of good throughput (75-100 tps) without losing much…

> It'll almost certainly be worth it, given the abusive behavior we've seen and will continue to see from the major closed-model providers. The proper financial comparison for GLM-5.2 would be one of the providers on OpenRouter or renting a server as needed. Compare apples to apples. You will almost certainly never break even compared to paying per token. Local LLMs at this scale are only worth it if you have extreme…

Never say never. When the free money party stops, then those token costs are going to have to go up and up. The fact there’s such a glaring disparity between the cost of running AI locally and the pennies it costs to use an online model shows how heavily funded those platforms are right now. This is not and cannot be sustainable.

Re: Jamesob's guide to running SOTA LLMs locally

#138

Earlier quoted context omitted.

Or if you want to hedge against the various tail risks of third-party providers raising prices or denying you service or somehow abusing your data...

> hedge against the various tail risks of third-party providers raising prices They could 10X the prices and you’d still be better off. It’s also unlikely that prices go up enough to warrant a $100K local investment to prevent paying a couple bucks per million tokens. > or denying you service I guess you’re not familiar with OpenRouter? There are many providers there. There are providers outside of OpenRouter. There…

People seem to miss that with local models you can have them burning their wee digital brains out 24/7, which is a different class of AI usage than that from online models even at a few dollars per million tokens.

Re: Jamesob's guide to running SOTA LLMs locally

#139
post #49

> "~$40k At this price level, you get the next step up in model intelligence. Something pretty close to Claude Opus." That is equivalent to 16.8 years of Claude Opus 4.8 or Codex GPT 5.5 at $200/mo. I'm a huge fan of running local models, but they're still wildly expensive, lower quality, and possibly dangerous (if backdoored). I sincerely wish this wasn't the case.

Stop trying to run them locally, folks. You don't own your fiber connection. So why try to own another rapidly depreciating, expensive, and annoying asset? Rent cloud GPUs! You get to participate in the ownership, data control, price control, and hacking culture without having to Frankenstein some hobbyist box that costs a ton, is distilled down to functional uselessness, and is a PITA to maintain.

This comment is like the antithesis of hacker culture. Can’t tell if being ironic or not.

Re: Jamesob's guide to running SOTA LLMs locally

#140

Earlier quoted context omitted.

> You don't own your fiber connection. So why try to own another rapidly depreciating, expensive, and annoying asset? Like a car? Because I don’t want to depend on uber or a taxi service. What does fiber have to do with anything? I don’t need the internet to run my local models.

Yes, "Like a car". LOL. You realise that many people in Europe and Asia do not own a car at all? Public transport, eBike / scooter, Tuk-Tuk, walk . The local LLM "privacy" war had been already lost.

You realise how many people in Europe do own a car?

And how exactly has the privacy war been lost?

Post reply on HN