Let's talk riser cables. I keep encountering issues with riser connectors claiming to support PCIe 4.0, which seem to have sub-par performance. They work fine with the GPUs and NICs I tested them with, but attaching a nvme drive causes all kinds of issues and prevents the machine from booting. I guess nvme isn't as tolerant of elevated bit-error-rates. That just doesn't inspire a lot of confidence in those risers, so…
All You Need Is 4x 4090 GPUs to Train Your Own Model
81–90 of 125 posts
Re: All You Need Is 4x 4090 GPUs to Train Your Own Model
#82Re: All You Need Is 4x 4090 GPUs to Train Your Own Model
#83Why not 3090s? Same VRAM and cheaper. With both setups you'd be limited to 1B. By contrast, you can run 4-bit quants of Llama 70B on two {3,4}090s, and it's still pretty lobotomized by modern standards. You can also train your own model even without GPUs. Just depends on parameter size.
It is previous architecture and it doesnt support newer version of Flash Attention , fp8 training etc
Re: All You Need Is 4x 4090 GPUs to Train Your Own Model
#84I would be much more intrested in a piece on what you can train with this kind of rig, rather than the rig itself
The bottleneck for most model training sizes is VRAM, and since each 4090 has 24 GB VRAM, that's 96 GB VRAM total. The article mentions that it can train LLMs from scratch up to 1 billion hyperparameters, which tracks. Nowadays that's not a lot: a single H100 that you can now rent has 80 GB VRAM, and doesn't have the technical overhead of handling work across GPUs.
Re: All You Need Is 4x 4090 GPUs to Train Your Own Model
#85Re: All You Need Is 4x 4090 GPUs to Train Your Own Model
#86Earlier quoted context omitted.
is your electricity free? Some of these cards probably cost about $0.10/hr to run ... depending on your card/electricity rate etc. It's probably somewhere between 12months-never depending on how the market shakes out. Maybe 2 years is a good idea ... really, if power is cheap/free and the machine is on and idle then it's free money - that's the way to look at it.
My electricity is not free, I would be satisfied with partially subsidizing these units too though
https://cloud.vast.ai/host/setup
There's a lot of competition in the "airbnb gpu" so if you don't like us, the number is around 12 or so globally. We're probably either #2 or #3. Companies don't really disclose these things so it's hard to know.
Some people probably list on more than one platform. There may be some host management software somewhere that helps with that. I haven't actually checked.
I'd be happy to talk more about these privately. Some are better than others and I've got no interest posting less than charitable things about our competitors publicly, regardless of how accurate I think it is. My email is in my profile.
Re: All You Need Is 4x 4090 GPUs to Train Your Own Model
#87Earlier quoted context omitted.
The bottleneck for most model training sizes is VRAM, and since each 4090 has 24 GB VRAM, that's 96 GB VRAM total. The article mentions that it can train LLMs from scratch up to 1 billion hyperparameters, which tracks. Nowadays that's not a lot: a single H100 that you can now rent has 80 GB VRAM, and doesn't have the technical overhead of handling work across GPUs.
Is there a reason you used hyperparameters rather than parameters? I was going to politely correct the terminology but you seem to be in AI for some time so either it was a mistype or I am misunderstanding what you are referencing.
Re: All You Need Is 4x 4090 GPUs to Train Your Own Model
#88Re: All You Need Is 4x 4090 GPUs to Train Your Own Model
#89You can get 4060 ti 16GB cards for ~$450 or 4070 ti 16gb for ~850 instead of the $2.5k for a 4090. I wonder how well 4 of those cards would perform. The 4060 TDP is 165w instead of 450w for the 4090. The 4070 looks like the best tradeoff though for cost/power/etc though. You could probably set up an 8 card 4070 ti 16gb system for less than the 4 card 4090 system
[1]: https://www.pugetsystems.com/labs/articles/llm-inference-con... (8GB model tested but it has same bus width and overall bandwidth as 16GB model)
[2]: https://www.reddit.com/r/LocalLLaMA/comments/1b5uwr4/some_gr...
[3]: https://www.reddit.com/r/LocalLLaMA/comments/178gkr0/perform...
Re: All You Need Is 4x 4090 GPUs to Train Your Own Model
#90Earlier quoted context omitted.
The last time I checked, a modern Threadripper build is a bit over $10,000. So if you have the budget for that but need something GPU-oriented instead, then I could see that being a reasonable option.
The thing is you need a threadripper-class build to make use of 4 GPUs in the first place. Ordinary PCs don't have the PCIe lanes necessary for that. But pricing is okay-ish, have a look at Geohot's Tinybox for turnkey solutions.