Live data from Hacker News

Benchmarking 15 “E-Waste” GPUs with Modern Workloads

esologic.com

41–50 of 68 posts

Re: Benchmarking 15 “E-Waste” GPUs with Modern Workloads

#41
post #7
post #6

Earlier quoted context omitted.

Interesting! I had only heard of them as cheap gaming boxes. Didn't know they were being used for cheap inference, too, but it makes sense. > They hold a special place in my heart because I deployed 20k Sounds like something I'd love to hear more about if you can share

ethereum mining, long shut down...

This needs a blog post & HN submission

Re: Benchmarking 15 “E-Waste” GPUs with Modern Workloads

#42
I’m obviously not the intended audience for this, and I understand this hardware is not useful for it, but I can’t help but feel an extra twinge of disappointment that there’s no mention of PC gaming anywhere in a post about GPUs in the comments here on HN. It says a lot.

Re: Benchmarking 15 “E-Waste” GPUs with Modern Workloads

#44

When I wanted to tinker with self-hosted models, I bought a couple of Radeon Pro V620 GPUs, because they're 32GB, still supported by current ROCm releases, and a few years newer than the similar-priced 32GB Nvidia cards (which are all EOL). They're a little faster than the old Tesla stuff, as well. 64GB is enough to run Gemma 4 31b 4-bit QAT with pretty big context at a respectable interactive speed (30+ tokens per s…

The B70 is woeful with respect to software performance today, unfortunately, and your stuck using Intels forks of things and it still doesn’t get the full expected throughput. Such a shame to be honest.

Re: Benchmarking 15 “E-Waste” GPUs with Modern Workloads

#45
post #44

When I wanted to tinker with self-hosted models, I bought a couple of Radeon Pro V620 GPUs, because they're 32GB, still supported by current ROCm releases, and a few years newer than the similar-priced 32GB Nvidia cards (which are all EOL). They're a little faster than the old Tesla stuff, as well. 64GB is enough to run Gemma 4 31b 4-bit QAT with pretty big context at a respectable interactive speed (30+ tokens per s…

The B70 is woeful with respect to software performance today, unfortunately, and your stuck using Intels forks of things and it still doesn’t get the full expected throughput. Such a shame to be honest.

Yeah, I'd recommend spending a little more for the AMD. As I understand, it's 40%+ faster. And, while ROCm is less mature than CUDA, it is miles more mature than the Intel stack.

Re: Benchmarking 15 “E-Waste” GPUs with Modern Workloads

#46

No mention of the venerable Tesla P4. 75W peak, 8GB VRAM, about $80 (£60). I have 6x P4s, a Xeon E5 2696v3 (36 threads, 3.8ghz peak but all core turbo unlocked, so 6 cores at 3.8Ghz - about 8 cores at 3.5ghz, or all cores at 3.1ghz), 48GB DDR4, all fit into a micro atx case running on a 650W MSI psu. This gives me a virtual 48GB GPU (llama.cpp ftw) to backup that 48GB of RAM. I typically see scores of at least 7-12t/…

That’s cool but 7 - 12 tps is frustrating for anything interactive.

Re: Benchmarking 15 “E-Waste” GPUs with Modern Workloads

#47

Earlier quoted context omitted.

> about $80 (£60) Man, I wish I lived where you guys lived.

Same, and hope I can afford that w/$ $/t, 650w is wild to run here24/7 , would cost me 32% of my salary

(650/1000)(24*30)=468 kWh.

That would cost me about $70/month ($0.15/kWh) or someone in California about $234/month ($0.50/kWh).

Do you pay $1/kWh and make ~$1500/month or something? I can’t make the math work for your case.

Re: Benchmarking 15 “E-Waste” GPUs with Modern Workloads

#48
Getting (further) into this myself so good timing. Running Qwen 3.6 27B at decent speed on some old cards but going to branch out.

I bough an Octominer for ~$150 which has power and PCIe slots and a basic Celeron and should let me expand to as many GPUs as I want.

I considered the P100s but I think the V100 16GBs are a better deal at $250. The 32GBs are way too much though.

Re: Benchmarking 15 “E-Waste” GPUs with Modern Workloads

#49
post #18

Earlier quoted context omitted.

"Don't post generated text" - https://news.ycombinator.com/newsguidelines.html

The rules for comments are now so lengthy that probably 90% of comments would violate one. It's like how police can pull you over for any reason and justify it by picking a law you've unwittingly violated

Easy; take your comment before you post it and have the LLM evaluate it against the rules.

Re: Benchmarking 15 “E-Waste” GPUs with Modern Workloads

#50

Would it possible to stack up to 16x32GB VRAM, and test the performance of a MOE model such as Deepseek-v4-flash?

16 GPUs would require one or more 220V breaker panels, more akin to an EV charger than a computer. You would also quickly run out of PCIe lanes. My goal with this benchmarking is to think about what is the most cost effective way to fill 4U.

Specifically, 16GPUs is extreme, but for what the GPUs are being used for, do they need all of the PCIe lanes since they are not pushing pixel data? I ask as someone I know was into extreme mods. He built a rig that he'd plug into his dryer's 220v outlet to run. He also had PCIe break out cables to plug in multiple GPUs per PCIe slot on the mobo. Since the GPUs were only doing math for 3D rendering, he did not worry about the PCIe lanes. This was a really long time ago and I do not remember the actual speeds, but faster than CPU only renders.
Post reply on HN