Live data from Hacker News

Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers

stability.ai

51–60 of 66 posts

Re: Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers

#51

Has anyone been running LLMs on TPUs in prod? Curious to hear experiences.

Yeah they train well and very stably even int8, maxtext now has LLaMA and mistral support too, pytorch xla gets 50% MFU with spmd and you have some nice stacks like levanter

Haven't been too impressed with inference versus tensor rt llm for example though

Re: Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers

#52
post #13

This is nice to foster some competition in hardware for model training, but the availability of these machines seems very limited - I don't think there's any major cloud provider allowing per hour rental of Gaudi2 VMs and Intel's own site directs you to buy an 8x GPU provisioned server from Supermicro for more than 40k USD. Availability and software stack is still heavily in Nvidia's favor right now, but maybe by the…

Disclaimer: I technically still am employed at Genesis Cloud (though no longer actively involved).

Genesis Cloud started integration and testing of Gaudi2 quite a while ago. I fully agree with the take of the article.

I can't promise per hour rental, but for longer times they are available! (should you be interested you can find contact details on the website)

Re: Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers

#53

Earlier quoted context omitted.

But you and I can't buy a TPU. You and I can buy an H100.

Speak for yourself! I can't even afford 1/10th of an H100.

He's probably speaking about availability, not affordability.

Re: Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers

#54

One question I have that nobody, including an Intel AXG employee, has been able to answer satisfsctorily for me is why both Gaudi and Ponte Vecchio exist. Wouldn't Intel have better chances of success if they focused on one product line?

From my understanding, Gaudi specializes in a specific use case (deep learning/AI) while Ponte Vecchio is more generic HPC. Also, DL/AI accelerators may not work well with newer models so the generic HPC hardware may be the only option for certain models until the DL/AI accelerators have a chance to catch up.

Re: Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers

#55
post #52
post #13

This is nice to foster some competition in hardware for model training, but the availability of these machines seems very limited - I don't think there's any major cloud provider allowing per hour rental of Gaudi2 VMs and Intel's own site directs you to buy an 8x GPU provisioned server from Supermicro for more than 40k USD. Availability and software stack is still heavily in Nvidia's favor right now, but maybe by the…

Disclaimer: I technically still am employed at Genesis Cloud (though no longer actively involved). Genesis Cloud started integration and testing of Gaudi2 quite a while ago. I fully agree with the take of the article. I can't promise per hour rental, but for longer times they are available! (should you be interested you can find contact details on the website)

Would you rent out a node for few days for benchmark testing?

Re: Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers

#56
post #13

This is nice to foster some competition in hardware for model training, but the availability of these machines seems very limited - I don't think there's any major cloud provider allowing per hour rental of Gaudi2 VMs and Intel's own site directs you to buy an 8x GPU provisioned server from Supermicro for more than 40k USD. Availability and software stack is still heavily in Nvidia's favor right now, but maybe by the…

Link for the price?

Re: Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers

#57

Earlier quoted context omitted.

I can feel the NVDA stock slipping as we speak… It has been amazing watching the groupthink at work on that stock when we just saw the same group do it on TSLA to disastrous effect. A similar no moat situation where they simply can’t imagine competitors ever existing.

The Model Y is the best selling car in the world in 2023. Those of us who were buying in 2019 are still up quite a bit even though the stock was higher at one point. RIVN, Ford, GM all are losing a lot of money on every EV they sell. We were right to bet on TSLA being a major winner. I actually put 40% of my TSLA into NVDA last year, because the demand for AI hardware is going to keep going up. I'm not saying the sto…

[deleted]

Re: Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers

#58
post #48

NVIDIA's profit margin is almost 92% on an H100. I'm surprised more chip companies haven't jumped on a "ML accelerator" bandwagon by now.

There's a dozen AI chips already; how many do you want? Now working ones is a different story.

For anyone deeply familiar with building one: what are the biggest problems you run into 3 months in that you didn't foresee?

Just curious because IME that's the point where the fun problems surface :)

Re: Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers

#59
post #58
post #48

Earlier quoted context omitted.

There's a dozen AI chips already; how many do you want? Now working ones is a different story.

For anyone deeply familiar with building one: what are the biggest problems you run into 3 months in that you didn't foresee? Just curious because IME that's the point where the fun problems surface :)

3 months in? I have not worked in chip design. But as an electronics engineer doing much less complex hardware (IoT sensors), I would say that is probably not enough time to hit any of the unexpected problems. Bringing an new ML accelerator architecture/family to market is likely 36 months at best.

Re: Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers

#60
post #15
post #13

This is nice to foster some competition in hardware for model training, but the availability of these machines seems very limited - I don't think there's any major cloud provider allowing per hour rental of Gaudi2 VMs and Intel's own site directs you to buy an 8x GPU provisioned server from Supermicro for more than 40k USD. Availability and software stack is still heavily in Nvidia's favor right now, but maybe by the…

Think it'll probably crack on with Gaudi3 at 4x performance, twice VRAM etc later this year. We found cuda sycl conversion surprisingly good https://www.intel.com/content/www/us/en/developer/articles/t...

It's hard to guess these cards' real performance uplifts. According to Nvidia, H100 is 11x faster than A100, but that's definitely not true in most cases. If Gaudi3 is legitimately 4x faster than Gaudi2, it should be a very good value proposition compared to even the B100. I'm really curious whether Intel will be able to compete with X100 using Falcon Shores or not. Regardless, I don't think Nvidia's margins are sustainable.
Post reply on HN