Live data from Hacker News

Google Launches AI Supercomputer Powered by Nvidia H100 GPUs

tomshardware.com

11–20 of 182 posts

Re: Google Launches AI Supercomputer Powered by Nvidia H100 GPUs

#11

So this is why Nvidia isn't lowering the price on the GPUs despite them sitting on the shelves and not selling. They make enough money from customers in the data center and supercomputer businesses that gaming is just a small market.

Gaming is a huge chunk of their revenue, around $2B in recent quarters, with datacenter around $3.5B.

Despite the AI hype, Nvidia’s datacenter revenue was down QoQ and only up 10% YoY.

It remains to be seen if the growth trajectory has changed meaningfully over the last quarter, because the stock is priced for massive earnings growth while their revenue and earnings have been actually shrinking.

We’ll find out on the upcoming earnings call

https://www.macrotrends.net/stocks/charts/NVDA/nvidia/revenu...

https://www.macrotrends.net/stocks/charts/NVDA/nvidia/eps-ea...

Re: Google Launches AI Supercomputer Powered by Nvidia H100 GPUs

#13
post #5

Does this mean google just deprecated TPUs? Not surprised.

this looks like it's for GCP. TPUs are used for most internal workloads. It's available externally but some of the papercuts and devex without the TPU/TF team helping you can be more painful than using Nvidia/CUDA

Re: Google Launches AI Supercomputer Powered by Nvidia H100 GPUs

#14
post #10
post #4

Earlier quoted context omitted.

250

Interesting, so what is the compute power of the 1000-node A100 super cluster my team has been allocated at work? I was expecting Google to be much bigger than us.

This is for Google Cloud users. My understanding is that Google mostly uses TPU internally.

Re: Google Launches AI Supercomputer Powered by Nvidia H100 GPUs

#16
post #15

Earlier quoted context omitted.

Why would you assume that?

Either GPUs are better for most AI tasks or TPUs. Both being overall approximately equally good is very unlikely.

> most AI tasks

Different workloads require different infrastructure.

Can your workload saturate the TPU without getting throttled by memory or network? Great! Use TPUs and reduce training cost.

But if your TPUs are idle 70% of the time because the constraint is getting data to them ...

"A3 represents the first production-level deployment of its GPU-to-GPU data interface, which allows for sharing data at 200 Gbps while bypassing the host CPU. This interface, which Google calls the Infrastructure Processing Unit (IPU), results in a 10x uplift in available network bandwidth for A3 virtual machines (VM) compared to A2 VMs."

Re: Google Launches AI Supercomputer Powered by Nvidia H100 GPUs

#17
post #5

Does this mean google just deprecated TPUs? Not surprised.

TPUs do compete with GPUs for ML tasks, so yes, this is evidence that GPUs are winning.

The only alternative I could imagine is that TPUs will "win" at supercomputers exclusivity aimed at inference (as opposed to training). Since TPUs excel at inference. The question is how much ML compute is used for inference as opposed to training. Not much, I guess, otherwise something like TPUs would be more popular.

Re: Google Launches AI Supercomputer Powered by Nvidia H100 GPUs

#18
post #17
post #5

Does this mean google just deprecated TPUs? Not surprised.

TPUs do compete with GPUs for ML tasks, so yes, this is evidence that GPUs are winning. The only alternative I could imagine is that TPUs will "win" at supercomputers exclusivity aimed at inference (as opposed to training). Since TPUs excel at inference. The question is how much ML compute is used for inference as opposed to training. Not much, I guess, otherwise something like TPUs would be more popular.

I've heard estimates that the amount of compute used to train GPT-4 is equivalent to 8 months of usage and most models are used much less than GPT-4 is, although I guess they are also easier to train.

Re: Google Launches AI Supercomputer Powered by Nvidia H100 GPUs

#19
Going Slightly Off Topic.

This is why Leading Edge Node will continue to be well funded. Consumer Electronics ( Mainly Smartphone ) Silicon usage has been the main push behind the development of Pure Play leading edge foundry in the past 10 years. Despite the predicted / expected drop of Smartphone sales, considering the potential shown by ChatGPT or Bard, GPU or Wafers dedicated for AI will continue to be in demand for at least another 5 years. In terms of lead time into the investment of silicon development that means we can continue to expect progress all the way till 2030, either 1nm or 0.8nm.

Re: Google Launches AI Supercomputer Powered by Nvidia H100 GPUs

#20
post #16
post #15

Earlier quoted context omitted.

Either GPUs are better for most AI tasks or TPUs. Both being overall approximately equally good is very unlikely.

> most AI tasks Different workloads require different infrastructure. Can your workload saturate the TPU without getting throttled by memory or network? Great! Use TPUs and reduce training cost. But if your TPUs are idle 70% of the time because the constraint is getting data to them ... "A3 represents the first production-level deployment of its GPU-to-GPU data interface, which allows for sharing data at 200 Gbps whi…

That's why I said "most" and "overall". Of course TPUs will have a niche. But it looks like the vast majority of money spent on ML compute is converging on GPUs.
Post reply on HN