Live data from Hacker News

Intel Gaudi 3 AI Accelerator

intel.com

41–50 of 260 posts

Re: Intel Gaudi 3 AI Accelerator

#41

Anyone have experience and suggestions for an AI accelerator? Think prototype consumer product with total cost preferably < $500, definitely less than $1000.

What else in on the BOM? Volume? At that price you likely want to use whatever resources are on the SoC that runs the thing and work around that. Feel free to e-mail me.

Re: Intel Gaudi 3 AI Accelerator

#43
post #37
post #25

Earlier quoted context omitted.

Does it not work for them? Where can I learn why?

Just go have a look around Github issues in their ROCm repositories on Github. A few months back the top excuse re: AMD was that we're not supposed to use their "consumer" cards, however the datacenter stuff is kosher. Well, guess what, we have purchased their datacenter card, MI50, and it's similarly screwed. Too many bugs in the kernel, kernel crashes, hangs, and the ROCm code is buggy / incomplete. When it works,…

[flagged]

Re: Intel Gaudi 3 AI Accelerator

#44
post #7

> Twenty-four 200 gigabit (Gb) Ethernet ports are integrated into every Intel Gaudi 3 accelerator WHAT‽ It's basically got the equivalent of a 24-port, 200-gigabit switch built into it. How does that make sense? Can you imaging stringing 24 Cat 8 cables between servers in a single rack? Wait: How do you even decide where those cables go? Do you buy 24 Gaudi 3 accelerators and run cables directly between every single…

The amount of power that will use up is massive, they should've gone for some fiber instead

The fiber optics are also extremely power hungry. For short runs people use direct attach copper cables to avoid having to deal with fiberoptics.

Re: Intel Gaudi 3 AI Accelerator

#45

Wow, I very much appreciate the use of the 5 Ws and H [1] in this announcement. Thank you Intel for not subjecting my eyes to corp BS [1] https://en.wikipedia.org/wiki/Five_Ws

I wonder if with the advent of LLMs being able to spit out perfect corpo-speak everyone will recenter to succint and short "here's the gist" as the long version will become associated to cheap automated output.

Re: Intel Gaudi 3 AI Accelerator

#46

Earlier quoted context omitted.

200gb is not going to be using CAT, it will be fiber (or direct attached copper cable as noted by dogma1138) with a QSFP interface

It will most likely use copper QSFP56 cables since these interfaces are either used in inter rack or adjacent rack direct attachments or to the nearest switch. O.5-1.5/2m copper cables are easily available and cheap and 4-8m (and even longer) is also possible with copper but tends to be more expensive and harder to get by. Even 800gb is possible with copper cables these days but you’ll end up spending just as much if…

Fair point!

Re: Intel Gaudi 3 AI Accelerator

#47
post #37

Earlier quoted context omitted.

Just go have a look around Github issues in their ROCm repositories on Github. A few months back the top excuse re: AMD was that we're not supposed to use their "consumer" cards, however the datacenter stuff is kosher. Well, guess what, we have purchased their datacenter card, MI50, and it's similarly screwed. Too many bugs in the kernel, kernel crashes, hangs, and the ROCm code is buggy / incomplete. When it works,…

[flagged]

Probably not, I have had the same experience with a 780m and a Mi60...

Re: Intel Gaudi 3 AI Accelerator

#48
post #37
post #25

Earlier quoted context omitted.

Does it not work for them? Where can I learn why?

Just go have a look around Github issues in their ROCm repositories on Github. A few months back the top excuse re: AMD was that we're not supposed to use their "consumer" cards, however the datacenter stuff is kosher. Well, guess what, we have purchased their datacenter card, MI50, and it's similarly screwed. Too many bugs in the kernel, kernel crashes, hangs, and the ROCm code is buggy / incomplete. When it works,…

It's seriously impressive how well AMD has been able to maintain their incredible software deficiency for over a decade now.

Re: Intel Gaudi 3 AI Accelerator

#49
post #37

Earlier quoted context omitted.

Just go have a look around Github issues in their ROCm repositories on Github. A few months back the top excuse re: AMD was that we're not supposed to use their "consumer" cards, however the datacenter stuff is kosher. Well, guess what, we have purchased their datacenter card, MI50, and it's similarly screwed. Too many bugs in the kernel, kernel crashes, hangs, and the ROCm code is buggy / incomplete. When it works,…

[flagged]

Effectively everyone that has attempted to use AMD hardware for ML comes away with these opinions, the main difference is how angrily they express it.

Re: Intel Gaudi 3 AI Accelerator

#50
post #15
post #8

Earlier quoted context omitted.

AMD MI300x is 192GB.

Which would be impressive had it _actually_ worked for ML workloads.

There's a number of scaled AMD deployments, including Lamini (https://www.lamini.ai/blog/lamini-amd-paving-the-road-to-gpu...) specifically for LLM's. There's also a number of HPC configurations, including the world's largest publicly disclosed supercomputer (Frontier) and Europe's largest supercomputer (LUMI) running on MI250x. Multiple teams have trained models on those HPC setups too.

Do you have any more evidence as to why these categorically don't work?

Post reply on HN