Live data from Hacker News

Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers

stability.ai

31–40 of 66 posts

Re: Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers

#31

Interesting! this was already the case with TPUs easily beating A100s. We sell Stable Diffusion finetuning on TPUs (dreamlook.ai), people are amazed how fast and cheap we can offer it - but there's no big secret, we just use hardware that's strictly faster and cheaper per unit of work. I expect a new wave of "your task, but on superior hardware" services to crop up with these chips!

Which TPUs do you use? Cloud-hosted or your own hardware? Interesting insight.

We started on V3s, now fully moved to V4s with some V5Es, investigating a full move towards V5E & V5P

Re: Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers

#32
post #2

"For Stable Diffusion 3, we measured the training throughput for the 2B Multimodal Diffusion Transformer (MMDiT) architecture model. Gaudi 2 trained images 1.5x faster than the H100-80GB, and 3x faster than A100-80GB GPU’s when scaled up to 32 nodes. "

I can feel the NVDA stock slipping as we speak… It has been amazing watching the groupthink at work on that stock when we just saw the same group do it on TSLA to disastrous effect. A similar no moat situation where they simply can’t imagine competitors ever existing.

It's a great company & will do well, plenty of demand & B100s/BH200s etc coming

The Hopper stuff is particulalry interesting

Re: Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers

#33

Gaudi3 is supposedly due this year with a 4X bump in Bf16 training over Gaudi2. Gaudi is an interesting product. Intel seems to have something pretty decent but it hasn't seen much of a volume release yet. Maybe that comes with V3? Not sure exactly what their strategy with it is. We do know that in 2025 it's supposed to be part of Intel's Falcon Shores HPC XPU. This essentially takes a whole bunch of HPC compute and…

It was interesting Aurora used GPU Max & definitely looking forward to Falcon Shores.

I think Gaudi2 was bad timed & they had to build stack, Gaudi3 is where I think we will see mass adoption given availability, way cheaper price/performance & maturer stack.

There is still weird stuff when using them but they are surprisingly solid.

Re: Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers

#35
post #2

"For Stable Diffusion 3, we measured the training throughput for the 2B Multimodal Diffusion Transformer (MMDiT) architecture model. Gaudi 2 trained images 1.5x faster than the H100-80GB, and 3x faster than A100-80GB GPU’s when scaled up to 32 nodes. "

I can feel the NVDA stock slipping as we speak… It has been amazing watching the groupthink at work on that stock when we just saw the same group do it on TSLA to disastrous effect. A similar no moat situation where they simply can’t imagine competitors ever existing.

The Model Y is the best selling car in the world in 2023. Those of us who were buying in 2019 are still up quite a bit even though the stock was higher at one point. RIVN, Ford, GM all are losing a lot of money on every EV they sell. We were right to bet on TSLA being a major winner.

I actually put 40% of my TSLA into NVDA last year, because the demand for AI hardware is going to keep going up. I'm not saying the stock will never go down, I'm sure it will be volatile, but don't confuse short term volatility with long term technologic transformations.

Re: Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers

#37

Interesting! this was already the case with TPUs easily beating A100s. We sell Stable Diffusion finetuning on TPUs (dreamlook.ai), people are amazed how fast and cheap we can offer it - but there's no big secret, we just use hardware that's strictly faster and cheaper per unit of work. I expect a new wave of "your task, but on superior hardware" services to crop up with these chips!

But you and I can't buy a TPU. You and I can buy an H100.

Re: Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers

#38
I'm wondering how AI scientists work these days. Do they really hack Cudakernels or do they plug models together with highlevel toolkits like pytorch?

Considering its the latter, considering pytorch takes care of providing optimized backends for various hardwares, how big of a moat is Cuda then really?

Re: Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers

#39
One question I have that nobody, including an Intel AXG employee, has been able to answer satisfsctorily for me is why both Gaudi and Ponte Vecchio exist. Wouldn't Intel have better chances of success if they focused on one product line?

Re: Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers

#40

One question I have that nobody, including an Intel AXG employee, has been able to answer satisfsctorily for me is why both Gaudi and Ponte Vecchio exist. Wouldn't Intel have better chances of success if they focused on one product line?

Gaudi was brought into Intel via an acquisition. Ponte Vecchio was an internal program. It can be explained by a combination of management silos and perhaps pre-existing obligations for Ponte Vecchio with the government for how they both came into being
Post reply on HN