Sorry to hijack the thread, but how would one get into managing GPU clusters? Modern GPUs are expensive, so it seems difficult to build a homelab to play around with them. Is learning how to run software on a cluster at the end-users level + playing around with VMs enough experience to enter the field?
> Modern GPUs are expensive, so it seems difficult to build a homelab to play around with them. Simulate a system with multiple high-end GPUs by setting up a system with one low-end GPU, breaking all the video outputs, and plugging it into a timeswitch that makes it lose power 3 times a week. Learn about industry norms by deciding it's crucial you have read access to production data, but at the same time that your us…
A practitioner's guide to testing and running GPU clusters
11–14 of 14 posts
Re: A practitioner's guide to testing and running GPU clusters
#12Glad to see the use of SLURM The number of times I see people trying to reinvent the HPC wheel astounds me
Sadly the fine article yada yadas the installation and integration of the GPUs with the scheduling of SLURM
Re: A practitioner's guide to testing and running GPU clusters
#13Sorry to hijack the thread, but how would one get into managing GPU clusters? Modern GPUs are expensive, so it seems difficult to build a homelab to play around with them. Is learning how to run software on a cluster at the end-users level + playing around with VMs enough experience to enter the field?
There isn't really a school for this stuff. The way I learned was to go work for a company that was building out GPU compute. It is a lot more than just software, especially on the high end of things.
- obligatory “Senior” in title
- requires 3-5 years of building physical GPU infrastructure
Re: A practitioner's guide to testing and running GPU clusters
#14Earlier quoted context omitted.
There isn't really a school for this stuff. The way I learned was to go work for a company that was building out GPU compute. It is a lot more than just software, especially on the high end of things.
Meanwhile in job seeking land: - obligatory “Senior” in title - requires 3-5 years of building physical GPU infrastructure
I never finished college. I also never let what was written job descriptions stop me from applying for a job that I wanted. I'm sure that the denials I have received have been all because of my own lack of performance during the interview.
I believe strongly in the whole "If there is a will, there is a way." If you get denied for one interview and you really want the job, you should try again at a later date. Find out what caused you to fail, fix it, and come back stronger.
I'm guessing that the number of people on the planet who've been hands-on in building large scale physical GPU infrastructure are in the low thousands. It isn't some huge field. We, as an industry, need people who really want to do this stuff, and can learn it quickly.
ProTip: you don't need experience with GPUs. I had zero when I started deploying 150,000 of them. What I had was an innate ability to figure shit out based on my other experiences, and that is what got me hired in the first place. I took ownership over the project and made it happen, no matter what it took. That's what people are looking for. Be a doer, not a talker.