Live data from Hacker News

Building an AI server on a budget

informationga.in

91–100 of 113 posts

Re: Building an AI server on a budget

#91
post #58

Good value but a 12GB card isn't going to let you do too much given the low quality of small models. Curious what "home AI" use cases small models are being used for? It would be nice to see a best value home AI setups under different budgets or RAM tiers, e.g. best value configuration for 128 GPU VRAM, etc. My 48GB GPU VRAM "Home AI Server" cost ~$3100 from all parts on eBay running 3x A4000's in a Supermicro 128GB…

I use a Proxmox server with RTX 3060 to generate paintings (I have a couple of old jailbroken Amazon Kindle's attached to walls for that purpose), and to run ollama, which is connected to Home Assistant & their voice preview device, allowing me to talk with LLM without transmitting anything to cloud services. Admittedly with that amount of VRAM the models I can run are fairly useless for stuff like controlling lights…

Why do you say that? You can easily finetune 8B parameter model for function calling.

Re: Building an AI server on a budget

#92

With system builds like this I always feel the VRAM is the limiting factor when it comes to what models you can run, and consumer grade stuff tends to max out at 16GB or (somemtimes) 24GB for more expensive models. It does make me wonder whether we'll start to see more and more computers with unified memory architecture (like the Mac) - I know nvidia have the Digits thing which has been renamed to something else

Go server GPU (TESLA) and 24 GB is not unusual. (And also about $300 used on eBay.)

But compute speed is very low.

Re: Building an AI server on a budget

#93
post #63
post #43

> You pay a lot upfront for the hardware, but if your usage of the GPU is heavy, then you save a lot of money in the long run. Last I saw data on this wasn’t true. A like for like comparison (same model and quant) API is cheaper than elec so you never make back hardware cost. That was a year ago and api costs have plummeted so I’d imagine it’s even worse now. Datacenters have cheaper elec, can do batch inference at s…

Is this also the case for token-heavy uses such as Claude Code? Not sure if I will end up using CC for development in the future, but if I end up leaning on that, I wonder if there would be a desire to essentially have it run 24/7. When ran 24/7, CC would possibly incur more API fees than residential electricity would cost when running on your own gear? I have no idea about the numbers. Just wondering.

I doubt you’re going to beat datacenter under any conditions in any model that is vaguely like for like

The comparison I saw was a small llama 8B model. ie something you can actually get usable numbers on both home and api. So something pretty commoditized

> When ran 24/7, CC would possibly incur more API fees than residential electricity would cost when running on your own gear?

Claude is pretty damn expensive so plausible that you can undercut it with another model. That implies you throw out the like for like assumptions out the door though. Valid play practically, but kinda undermines the buy own rig to save argument

Re: Building an AI server on a budget

#94

>DECISION: Nvidia RTX 4070 I'm curiuos why OP didn't go for the more recent Nvidia RTX 4060 Ti with 16 GB VRAM that cost cheaper (~USD500) brand new and lesser power consumption at 165W [1]. [1] RTX 5060 Ti 16GB sucks for gaming, but seems like a diamond in the rough for AI: https://news.ycombinator.com/item?id=44196991

And if you're gonna be fine with 12GB, why not a 2080ti instead?

Only 11GB... but I guess it will allow you to not do anything useful just as well as 12GB will :)

You can however solder on double-capacity memory chips to get 22GB:

https://forums.overclockers.com.au/threads/double-your-gpu-m...

I hoped the article would be more along these lines than calling an unremarkable second-hand last-gen gaming pc an "AI Server".

Re: Building an AI server on a budget

#97

With system builds like this I always feel the VRAM is the limiting factor when it comes to what models you can run, and consumer grade stuff tends to max out at 16GB or (somemtimes) 24GB for more expensive models. It does make me wonder whether we'll start to see more and more computers with unified memory architecture (like the Mac) - I know nvidia have the Digits thing which has been renamed to something else

That’s what I hope for, but everything that isn’t bananas expensive with unified memory has very low memory bandwidth. DGX (Digits), Framework Desktop, and non-Ultra Macs are all around 128 gb/s, and will produce single digits tokens per second for larger models: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inferen...

So there’s a fundamental tradeoff between cost, inference speed, and hostable model size for the foreseeable future.

Re: Building an AI server on a budget

#98

What are the practical uses of a self hosted LLM? Is it actually possible to approach the likes of Claude or one of the other big ones on your own hardware for a reasonable budget? I don’t know if this is something that’s actually worth it or if people are just building these rigs for fun or niche use cases that don’t require the intelligence of a hosted LLM.

Personal opinion, it's for fun with some internal narrative of justification. It doesn't seem like it would be cost effective or provide better results, as all the major LLM vendors benefit tremendously from economies of scale, and the monthly fees for these services are extremely reasonable for what you are getting. Going further, the cloud based LLM receive upgrades constantly while static hardware will likely lock you out of future models at some time horizon.

Re: Building an AI server on a budget

#99

Whenever I get to a section that was clearly autogenerated by an LLM I lose interest in the entire article. Suddenly the entire thing is suspect and I feel like I’m wasting my time, since I’m lo lingering encountering the mind of another person, just interacting with a system.

Eh, yeah - the article starts off pretty specific but then gets into the weeds of stuff like how to put your PC together, which is far from novel information and certainly not on-topic in my opinion.

I sent the article link to my son because he does not have experience building or assembling hardware or installing or using Linux. Also took the author's ChatGPT prompt and changed it to ask about reusing two HPE ML150 Gen9 servers I picked up free. I think my son will benefit from the details in the article that many find off-topic.

Re: Building an AI server on a budget

#100

Someone posted that they had used a "mining rig" [0] from AliExpress for less than $100. It even has RAM and a CPU. He picked up a 2000W (!) DELL server PS for cheap off eBay. The GPUs were NVIDIA TESLAs (M40 for example) since they often have a lot of RAM and are less expensive. I followed in those footsteps to create my own [1] (photo [2]). I picked up a 24GB M40 for around $300 off eBay. I 3D printed a "cowl" for…

My first guess would be to change the Above 4G decoding setting but depending upon how old the motherboard is it may not have that setting.
Post reply on HN