Live data from Hacker News

Documentation for the AMD 7900XTX

github.com

121–130 of 154 posts

Re: Documentation for the AMD 7900XTX

#121
post #44

I used to be an AMD diehard fan. Bought their processors and GPUs for close to 20 years until I simply couldn't take it anymore - 5 years ago I started buying NVIDIA GPUs with Intel CPUs and honestly never looked back. Sure I pay more, but it's worth paying more for reliability and software support. Anything else and I'm losing money out of ideology to give to a business that doesn't have its act together.

Funny, AMD CPUs have started becoming viable (and then better than Intel) five years ago.

Re: Documentation for the AMD 7900XTX

#122

Earlier quoted context omitted.

> The problem is if I spend an evening trying to do anything with OpenCL or ROCm the kernel hard-locks and I go to bed early. In a race against NVIDIA where every comment on here is about just how much better CUDA is than any alternative, why doesn't literally what you just said "also" have the attention of Lisa Su?

I have had multiple replies from Lisa Su on both Twitter and by e-mail. It doesn't help. AMD is structurally incapable of fixing these issues. They don't even have a 7900XTX in CI. Crashing the firmware is so trivial there's no way anyone fuzzed anything. They debug by the application, adding mitigations at many layers of the stack to make it work. While this strategy is fine for the 20 mainstream games that come out…

Yikes. For context, the 7900XTX and 7900XT are the only consumer GPUs that AMD officially support running AI workloads on, and the XTX is the one they recommend out of the two. Other cards in the same generation and any from previous generations are officially not supported or tested and may not work. So it's not like they have a particularly large amount of hardware that needs to be included in their testing environment.

Re: Documentation for the AMD 7900XTX

#123
post #8

So for reference, this geohot is George Hotz who has a company Tiny Corp [0] in the space. Among a long list of things, he had a moment in the spotlight recently for a long rant [1] where he "gave up on AMD" because their drivers sucked more than he expected. The situation is reasonably complex - AMD have some programmers working on ROCm who seemed to be operating at standard fare when that is really not what AMD nee…

I’m going remind everyone that AMD and NVIDIA are atypically competition avoidant, but I don’t want to defend a generalization that broadcast let’s be specific: Hopper and MI300 are on paper just straight up competitive, both super scarce, high-margin cards that really love PyTorch, (which really loves NVIDIA, even TPU is, not the fist among equals). But all the MI300 is going to supercomputer and adjacent things, an…

>I’m going remind everyone that AMD and NVIDIA are atypically competition avoidant, but I don’t want to defend a generalization that broadcast let’s be specific: Hopper and MI300 are on paper just straight up competitive, both super scarce, high-margin cards that really love PyTorch, (which really loves NVIDIA, even TPU is, not the fist among equals). > But all the MI300 is going to supercomputer and adjacent things, and all the Hopper is going to giant tech LLM type stuff, it’s not the same people bidding on those bins of those.

They clearly compete fiercely, every feature NVIDIA adds to its consumer GPUs gets an equivalent version by AMD. Previously AMD has introduced features that have caught on enough to force NVIDIA to implement them. AMD's offerings often force NVIDIA to discontinue certain products or to drop prices. The reason MI300 is going to supercomputing and Hopper to AI is the self fulfilling prophecy that AMD's software still sucks for AI, thus they have to target general supercomputing.

>Oh, and their respective CEOs are closely blood-related and in a trivially first name if not family gathering basis.

This is mostly just a meme, they're distant relatives, which is not too surprising considering they're both from Taiwan. Jensen's maternal uncle is Lisa's grandfather, with said maternal uncle being the eldest of 12 siblings, 18 years older than Jensen's mother. I barely know the names of the eldest of my aunts, who have a similar age gap, let alone have a close relationship with them, because that age gap and number of siblings means that even my mother barely knows them.

Re: Documentation for the AMD 7900XTX

#124

Earlier quoted context omitted.

NVidia is making boatloads of money because their driver works and they have a software library called CUDA that accelerates neural networks. Nobody expects AMD to match them, but George thought they could at least write a GPU driver. If AMD can get that driver out, then George could provide a competitor to CUDA (for neutral networks only). They'd both make boatloads of money. However, AMD was less capable than expec…

>However, AMD was less capable than expected and their drivers were too buggy to run neural networks like those needed for the MLPerf benchmark. So now, it appears that AMD, Tinybox, and investors like me won't be making boatloads of money. This is where the melodrama kicks in. They reverse-reversed course less than a week later and now they're back on AMD again. AMD plans go "on hold" - March 19: https://twitter.com…

> This is where the melodrama kicks in.

The complaint was that they couldn't debug the GPU.

Now they can, so they're soldiering on.

Seems reasonable to me.

Re: Documentation for the AMD 7900XTX

#125

Earlier quoted context omitted.

I’m going remind everyone that AMD and NVIDIA are atypically competition avoidant, but I don’t want to defend a generalization that broadcast let’s be specific: Hopper and MI300 are on paper just straight up competitive, both super scarce, high-margin cards that really love PyTorch, (which really loves NVIDIA, even TPU is, not the fist among equals). But all the MI300 is going to supercomputer and adjacent things, an…

>I’m going remind everyone that AMD and NVIDIA are atypically competition avoidant, but I don’t want to defend a generalization that broadcast let’s be specific: Hopper and MI300 are on paper just straight up competitive, both super scarce, high-margin cards that really love PyTorch, (which really loves NVIDIA, even TPU is, not the fist among equals). > But all the MI300 is going to supercomputer and adjacent things,…

This is a contentious debate of which I merely think everyone interested in this should be aware. I included my position, which is shared by many, that this is ridiculous if not illegal.

I hope I didn't confuse anyone as to there being a different and dramatically better funded side of the argument, given that everyone on HN is squarely the target audience for the PR blitz on this. Everyone knows all the arguments for why we're all supposed to say "This is fine. Everything is great here.".

A much milder position than "cousins are running companies that seem to be coordinating" is: "55.58% Net Profit Margin last quarter isn't consistent with a functioning market". Those are Wintel margins, those are "get more than noticed by the DoJ" margins.

Throw in everything `ROCm` down to the "hangs randomly doing ostensibly supported things on a mainstream platform like Ubuntu 22.04 with a modern kernel" being pretty clearly under-resourced for an organization that can do Zen4 and it's basically flawless firmware / UEFI / driver / etc. story?

AMD can build a 7900XT that is a joy to use for people who don't need 100Gb of VRAM or FP8 training. They'd sell a zillion of them at margins that even EPYC would envy, and they've had people on their ass in public about it for coming up on a few years now.

NVIDIA can build a prosumer card, and continue to be the GPGPU vendor that all of us had no compunctions about listing as one of our favorite companies in tech (with a few exceptions like the Linux fiasco ages ago) even a couple of years ago. The "holy shit you can do all this on a 3090-Ti and a normal tech worker can save up enough to buy one pretty easily" days were imperfect, but I still really liked NVIDIA, or at least enough not to be desperately looking for other options as my default posture.

Both companies can easily make a pile, look like the good guys, draw hackers into the ecosystem, and enjoy both the wads of cash and love that everyone would be throwing at them. The competition would be better for both teams! It just wouldn't turn into a short-term asset, it would turn into a long-term asset: it would be "our team is in true fighting form on this", and you can't report that next quarter. And the parts of AMD and NVIDIA that have serious pressure on them, that are in fighting shape? Those teams are red hot, they're killing it.

They just don't want to.

Re: Documentation for the AMD 7900XTX

#126
post #8

So for reference, this geohot is George Hotz who has a company Tiny Corp [0] in the space. Among a long list of things, he had a moment in the spotlight recently for a long rant [1] where he "gave up on AMD" because their drivers sucked more than he expected. The situation is reasonably complex - AMD have some programmers working on ROCm who seemed to be operating at standard fare when that is really not what AMD nee…

I’m going remind everyone that AMD and NVIDIA are atypically competition avoidant, but I don’t want to defend a generalization that broadcast let’s be specific: Hopper and MI300 are on paper just straight up competitive, both super scarce, high-margin cards that really love PyTorch, (which really loves NVIDIA, even TPU is, not the fist among equals). But all the MI300 is going to supercomputer and adjacent things, an…

> But all the MI300 is going to supercomputer

I'm building a bare metal cloud service provider entirely around MI300x. Anyone (within US export restrictions) can have reasonably priced access to them. You get full access to the cards (not API based). If you take an entire machine (of cluster of them), you get BMI level access. It is as if you're sitting at the machine yourself.

In other words, I'm building my own supercomputer, and making it public.

Re: Documentation for the AMD 7900XTX

#127

Earlier quoted context omitted.

I’m going remind everyone that AMD and NVIDIA are atypically competition avoidant, but I don’t want to defend a generalization that broadcast let’s be specific: Hopper and MI300 are on paper just straight up competitive, both super scarce, high-margin cards that really love PyTorch, (which really loves NVIDIA, even TPU is, not the fist among equals). But all the MI300 is going to supercomputer and adjacent things, an…

> But all the MI300 is going to supercomputer I'm building a bare metal cloud service provider entirely around MI300x. Anyone (within US export restrictions) can have reasonably priced access to them. You get full access to the cards (not API based). If you take an entire machine (of cluster of them), you get BMI level access. It is as if you're sitting at the machine yourself. In other words, I'm building my own sup…

Thank you for two things: first thank you for correcting my error is saying "all the MI300" is going to supercompute, I am aware that a few groups are doing awesome stuff with making the platform available to mortals.

Second, thank you for being one of the people trying to make the platform available to mortals. As Sheryl used to say when someone did the right thing: "You're doing God's work."

Is there a link or a mailinglist or something that I can watch for how to get access when it becomes available? I'd like to have `HYPER // MODERN // AI` support the platform well.

edit: It's in your profile. I will check it out tonight. Keep it up!

Re: Documentation for the AMD 7900XTX

#128

Earlier quoted context omitted.

I thought he sort of gave up and is going to build his stuff on Nvidia in the end? It's wild. I'd want AMD to present some sort of alternative if only because you can't get any high-mem chips these days but they seem to have some massive organizational uselessness on this front.

They changed their mind a few days later and announced they're going to do NVIDIA and AMD versions of their box.

Interesting news. Thank you for that update.

Re: Documentation for the AMD 7900XTX

#129

Earlier quoted context omitted.

>However, AMD was less capable than expected and their drivers were too buggy to run neural networks like those needed for the MLPerf benchmark. So now, it appears that AMD, Tinybox, and investors like me won't be making boatloads of money. This is where the melodrama kicks in. They reverse-reversed course less than a week later and now they're back on AMD again. AMD plans go "on hold" - March 19: https://twitter.com…

> This is where the melodrama kicks in. The complaint was that they couldn't debug the GPU. Now they can, so they're soldiering on. Seems reasonable to me.

Surely if they had done a good job communicating with AMD they could have figured that out beforehand. The repo was already public.

Re: Documentation for the AMD 7900XTX

#130

Earlier quoted context omitted.

> But all the MI300 is going to supercomputer I'm building a bare metal cloud service provider entirely around MI300x. Anyone (within US export restrictions) can have reasonably priced access to them. You get full access to the cards (not API based). If you take an entire machine (of cluster of them), you get BMI level access. It is as if you're sitting at the machine yourself. In other words, I'm building my own sup…

Thank you for two things: first thank you for correcting my error is saying "all the MI300" is going to supercompute, I am aware that a few groups are doing awesome stuff with making the platform available to mortals. Second, thank you for being one of the people trying to make the platform available to mortals. As Sheryl used to say when someone did the right thing: "You're doing God's work." Is there a link or a ma…

Thanks!

We started the business last year, as a proof of concept before MI300x was even released. After a lot of work to build the necessary relationships, we got a box of MI300x, deployed it into our data center, and immediately onboarded a very large customer.

Now that we've completed that PoC phase of the business, we just closed many millions in additional funding that will go purchase many more GPUs. Thanks to our hard work and excellent investors, it is actually happening!

While we wait for more GPUs, we are donating time on the box to anyone who'd like to run benchmarks [0] and publish unbiased blog posts (with repeatable source code). Our hope is that once people see how well these perform, they will consider porting/running their code on our systems.

If they don't perform we will be transparent about that as well. I'll have a good set of data to bring back to AMD/SMCI to try to resolve any issues. My guess is that AMD will take that a lot more seriously than someone trying to make consumer hardware work, in an enterprise setting. This isn't a knock on George's ambitions, nor the need for AMD to make consumer products work better with ROCm/AI. I just feel that AMD has limited resources today and they have to focus on one thing first, which is what they are clearly doing in their responses to him.

[0] https://www.reddit.com/r/LocalLLaMA/comments/1bpgrdf/wanted_...

Post reply on HN