Live data from Hacker News

AMD may get across the CUDA moat

hpcwire.com

91–100 of 312 posts

Re: AMD may get across the CUDA moat

#91

Earlier quoted context omitted.

I ran 150,000+ AMD cards for mining ETH. Once I fully automated all the vbios installs and individual card tuning, it ran beautifully. Took a lot of work to get there though! Fact is that every single GPU chip is a snowflake. No two operate the same.

Have you ever written about this enterprise? This sounds super unique and I would be very interested in hearing about how it was run and how it turned out.

It was unique, not many people on the planet, that I know of, who've run as many GPUs as I have. Especially not working for a giant company with large teams of people. For the tech team, it was just me and one other guy. Everything had to be automated because there was no way we could survive otherwise.

I've put a bunch of comments here on HN about the stuff I can talk about.

It no longer exists after PoS.

Re: AMD may get across the CUDA moat

#92
post #42
post #26

CUDA is the only reason I have an Nvidia card, but if more projects start migrating to a more agnostic environment, I'll be really grateful. Running Nvidia in Linux isn't as much fun. Fedora and Debian can be incredibly reliable systems, but when you add an Nvidia card, I feel like I am back in Windows Vista with kernel crashes from time to time.

Yeah, nvidia linux support is meh, but still much better than amd.

Not my experience. The open source AMD drivers are much more pleasant to deal with than the closed source Nvidia ones.

Re: AMD may get across the CUDA moat

#94
post #26

CUDA is the only reason I have an Nvidia card, but if more projects start migrating to a more agnostic environment, I'll be really grateful. Running Nvidia in Linux isn't as much fun. Fedora and Debian can be incredibly reliable systems, but when you add an Nvidia card, I feel like I am back in Windows Vista with kernel crashes from time to time.

I see these complains from time to time and I never understand them. I've literally been running nvidia on linux since the TNT2 days and have _never_ had this sort of issue. That's across many drivers and many cards over the many many years.

I understand it, but I also haven't had any trouble since I figured out the right procedure for me on fedora (which probably took some time, but it's been so long that I can't remember). Whenever I read people having issues it sounds like they are using a package installed via dnf for the driver/etc. I've always had issues with dkms and the like and just install the latest .run from nvidia's website whenever I have a kernel update (I made a one-line script to call it with the silent option and flags for signing for secure boot so I don't really think about it). No issues in a very long time even with the whackiness of prime/optimus offloading on my old laptop.

Re: AMD may get across the CUDA moat

#95
post #46
post #42

Earlier quoted context omitted.

Yeah, nvidia linux support is meh, but still much better than amd.

Is it better than AMD? I have had literally no graphics issues on my 6650 XT with swaywm using the built in kernel drivers.

This week I upgraded my kernel on a 2017 workstation to 6.5.5 and when I rebooted and looked at 'dmesg' there were no less than 7 kernel faults with stack traces in my 'dmesg' from amdgpu. Just from booting up. This is a no-graphical-desktop system using a Radeon Pro W5500, which is 3.5 years old (I just had the card and needed something to plug in for it to POST.)

I have come to accept that graphics card drivers and hardware stability ultimately comes down to whether or not ghosts have decided to haunt you.

Re: AMD may get across the CUDA moat

#96

Earlier quoted context omitted.

Have you ever written about this enterprise? This sounds super unique and I would be very interested in hearing about how it was run and how it turned out.

It was unique, not many people on the planet, that I know of, who've run as many GPUs as I have. Especially not working for a giant company with large teams of people. For the tech team, it was just me and one other guy. Everything had to be automated because there was no way we could survive otherwise. I've put a bunch of comments here on HN about the stuff I can talk about. It no longer exists after PoS.

what type of cards did you have? what did you do with them after PoS? How did you even buy so many cards? Sorry, like the other commenter I'm extremely curious

Re: AMD may get across the CUDA moat

#97
post #88
post #72

There is only limited empirical evidence of AMD closing the gap that NVidia has created in the science or ML software. Even when considering pytorch only, the engineering effort to maintain specialized ROCm along with CUDA solutions is not trivial (think flashattention, or any customization that optimizes your own model). If your GPUs only need a simple ML workflow all times for a few years nonstop, maybe there exist…

The fact that El Capitan is AMD says that at least for Science/HPC there definitely is evidence of a closing gap.

Thanks. You are actually right that this new supercomputer might move the needle once it is in production mode. I will wait and see how it goes.

Re: AMD may get across the CUDA moat

#98

Earlier quoted context omitted.

Didn't he do what he always does. Rake in a ton of money, fart around and then cash out exclaiming it's everyone else's fault? The way he stole Fail0verflow's work with the PS3 security leak after failing to find a hypervisor exploit for months absolutely soured any respect I had for him at the time

Yep, did exactly that. IMO he threw a fit, even though AMD was working with him squashing bugs. https://github.com/RadeonOpenCompute/ROCm/issues/2198#issuec...

He's back on it after getting AMD's CEO to commit resources to this:

https://twitter.com/realGeorgeHotz/status/166980346408248934...

https://twitter.com/LisaSu/status/1669848494637735936

Re: AMD may get across the CUDA moat

#99
post #40

Earlier quoted context omitted.

I think in this case the changes needed to make AMD useful will open the market to other players as well (e.g. Intel). PyTorch is already walking down this path and while CUDA-based performance is significantly better, that is changing and of course an area of continued focus. It's not that people don't like Nvidia, rather it's just that there is a lot of hardware out there that can technically perform competitively,…

Last I checked I saw the H100 was about two gens more advanced for certain components (tensor cores, bfloats, cache, mem bandwidth) - but my research may have been wrong as admittedly I'm not as familiar with AMDs offerings for GPU.

They are not behind... https://www.tomshardware.com/news/amd-expands-mi300-with-gpu...

You can also actually buy them as opposed to the nVidia offerings which you are going to have to fight for.

Re: AMD may get across the CUDA moat

#100

Earlier quoted context omitted.

It was unique, not many people on the planet, that I know of, who've run as many GPUs as I have. Especially not working for a giant company with large teams of people. For the tech team, it was just me and one other guy. Everything had to be automated because there was no way we could survive otherwise. I've put a bunch of comments here on HN about the stuff I can talk about. It no longer exists after PoS.

what type of cards did you have? what did you do with them after PoS? How did you even buy so many cards? Sorry, like the other commenter I'm extremely curious

Primarily 470,480,570,580. We also ran a very large cluster of PS5 APU chips too.

Got the chips directly from AMD. Since these are 4-5 year old chips, they were not going to ever be used. It is more ROI efficient with ETH mining to use older cards than newer ones.

Had a couple OEM manufacture the cards specially for us with 8gb, heatsinks instead of fans (lower power usage) and no display ports (lower cost).

They will be recycled as there isn't much use for them now.

I'm also no longer with the company.

Post reply on HN