Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers
1–10 of 66 posts
Re: Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers
#2Re: Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers
#3"For Stable Diffusion 3, we measured the training throughput for the 2B Multimodal Diffusion Transformer (MMDiT) architecture model. Gaudi 2 trained images 1.5x faster than the H100-80GB, and 3x faster than A100-80GB GPU’s when scaled up to 32 nodes. "
It has been amazing watching the groupthink at work on that stock when we just saw the same group do it on TSLA to disastrous effect. A similar no moat situation where they simply can’t imagine competitors ever existing.
Re: Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers
#4I expect a new wave of "your task, but on superior hardware" services to crop up with these chips!
Re: Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers
#5[1] https://www.intel.com/content/www/us/en/developer/articles/t...
Re: Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers
#6"For Stable Diffusion 3, we measured the training throughput for the 2B Multimodal Diffusion Transformer (MMDiT) architecture model. Gaudi 2 trained images 1.5x faster than the H100-80GB, and 3x faster than A100-80GB GPU’s when scaled up to 32 nodes. "
I can feel the NVDA stock slipping as we speak… It has been amazing watching the groupthink at work on that stock when we just saw the same group do it on TSLA to disastrous effect. A similar no moat situation where they simply can’t imagine competitors ever existing.
* just to be clear - this is a joke
Re: Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers
#7Re: Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers
#8Re: Intel Gaudi2 chips outperform Nvidia H100 on diffusion transformers
#9"For Stable Diffusion 3, we measured the training throughput for the 2B Multimodal Diffusion Transformer (MMDiT) architecture model. Gaudi 2 trained images 1.5x faster than the H100-80GB, and 3x faster than A100-80GB GPU’s when scaled up to 32 nodes. "