Live data from Hacker News

Google Launches AI Supercomputer Powered by Nvidia H100 GPUs

tomshardware.com

51–60 of 182 posts

Re: Google Launches AI Supercomputer Powered by Nvidia H100 GPUs

#51
post #35

I'm more interested in what normal folks are running at home. What are your builds?

Honestly it's a $79 Lonovo 3 Chromebook running a Gcloud A3 virtual workstation over 5G from the golf course ;)

Whats the battery life on that :)

Re: Google Launches AI Supercomputer Powered by Nvidia H100 GPUs

#52
post #16
post #15

Earlier quoted context omitted.

Either GPUs are better for most AI tasks or TPUs. Both being overall approximately equally good is very unlikely.

> most AI tasks Different workloads require different infrastructure. Can your workload saturate the TPU without getting throttled by memory or network? Great! Use TPUs and reduce training cost. But if your TPUs are idle 70% of the time because the constraint is getting data to them ... "A3 represents the first production-level deployment of its GPU-to-GPU data interface, which allows for sharing data at 200 Gbps whi…

TPU's have a TPU-TPU interconnect that is faster and lower latency than any GPU cluster [1]. That said this is a huge leap for GPU's on GCP. For A100's SOTA is 1.6tbit per host over Infiniband (which azure and some smaller gpu clouds provided), AWS had 400-800 Gbit and GCP had .... ~100gbit.

SOTA seems to be 3.2Tbit for H100 clusters so this still seems a bit slow? (Tricky as they don't give us a clear number just 10x). H100's are much more powerful per chip though so at least initially the clusters will be smaller and not network bound.

The tricky thing is no one other than Azure of the big providers seems willing to pay Nvidia's margins for RDMA switches, it seems this is still the case.

[1] https://arxiv.org/pdf/2304.01433.pdf

Re: Google Launches AI Supercomputer Powered by Nvidia H100 GPUs

#53
post #47

Does this mean Google is giving up on TPUs? TPUs were supposed to be their unfair advantage in the cloud ML/DL space. But from what I've experienced, and have heard from other engineers, there's always some subtle incompatibility with TPUs that requires modifying the training/eval scripts. I wonder why they didn't try to polish the rough edges with Pytorch, et al. If they're admitting TPUs aren't their competitive ad…

There's a few comments to this effect in the thread, and I don't entirely understand where they're coming from. There's nothing in the article suggesting they've changed their strategy with TPUs in any way. The word TPU isn't even mentioned here. There's no suggestion they're actually using this internally either. There's no benchmarks showing that it's more cost-effective or scales better. And isn't your second para…

>There's nothing in the article suggesting they've changed their strategy with TPUs in any way

Google owns and designs their own TPUs. They offer these TPUs in the cloud. I've seen many comments in here about how next-level TPUs are (despite zero evidence indicating that). Google even disclaims their TPU by saying that you shouldn't compare it with the H100 given node levels et al.

Their premiere offering is an nvidia H100 offering.

Yes, of course this is a pretty telling indication. If Google was all in on TPUs they'd be building mega TPU systems and pushing those. Instead they're pushing nvidia AI offerings.

Re: Google Launches AI Supercomputer Powered by Nvidia H100 GPUs

#55

Does this mean Google is giving up on TPUs? TPUs were supposed to be their unfair advantage in the cloud ML/DL space. But from what I've experienced, and have heard from other engineers, there's always some subtle incompatibility with TPUs that requires modifying the training/eval scripts. I wonder why they didn't try to polish the rough edges with Pytorch, et al. If they're admitting TPUs aren't their competitive ad…

This is for GCP. Google themselves probably still trains on custom hardware but they don't offer their latest and greatest hardware on GCP.

Offering more options to customers is always better especially when Nvidia has great market share in this area. This is probably the reason why Microsoft is trying to help AMD catch up so their is more competition. AI GPU prices are insane compared to standard GPU because of the lack of competition.

Re: Google Launches AI Supercomputer Powered by Nvidia H100 GPUs

#56
post #47

Does this mean Google is giving up on TPUs? TPUs were supposed to be their unfair advantage in the cloud ML/DL space. But from what I've experienced, and have heard from other engineers, there's always some subtle incompatibility with TPUs that requires modifying the training/eval scripts. I wonder why they didn't try to polish the rough edges with Pytorch, et al. If they're admitting TPUs aren't their competitive ad…

There's a few comments to this effect in the thread, and I don't entirely understand where they're coming from. There's nothing in the article suggesting they've changed their strategy with TPUs in any way. The word TPU isn't even mentioned here. There's no suggestion they're actually using this internally either. There's no benchmarks showing that it's more cost-effective or scales better. And isn't your second para…

The word TPU isn't even mentioned here.

A sentence that reads "I am going to eat nothing but vegetables from now on" doesn't mention meat, but you can infer that I won't eat meat again from the sentence.

A sentence that says Google are going all in on nVidea GPUs for AI doesn't need to mention TPUs to convey information about their future either.

Re: Google Launches AI Supercomputer Powered by Nvidia H100 GPUs

#57

Does this mean Google is giving up on TPUs? TPUs were supposed to be their unfair advantage in the cloud ML/DL space. But from what I've experienced, and have heard from other engineers, there's always some subtle incompatibility with TPUs that requires modifying the training/eval scripts. I wonder why they didn't try to polish the rough edges with Pytorch, et al. If they're admitting TPUs aren't their competitive ad…

If they are selling gpu compute, nobody wants to use a Google TPU, they want cuda

Re: Google Launches AI Supercomputer Powered by Nvidia H100 GPUs

#59
post #57

Does this mean Google is giving up on TPUs? TPUs were supposed to be their unfair advantage in the cloud ML/DL space. But from what I've experienced, and have heard from other engineers, there's always some subtle incompatibility with TPUs that requires modifying the training/eval scripts. I wonder why they didn't try to polish the rough edges with Pytorch, et al. If they're admitting TPUs aren't their competitive ad…

If they are selling gpu compute, nobody wants to use a Google TPU, they want cuda

And they want to support people migrating from other cloud providers where they are already using nvidia/Cuda. Though it also helps support the opposite migration, they are the smaller cloud player trying to get customers, not the big one trying to constrain them as much yet.

Re: Google Launches AI Supercomputer Powered by Nvidia H100 GPUs

#60
post #56
post #47

Earlier quoted context omitted.

There's a few comments to this effect in the thread, and I don't entirely understand where they're coming from. There's nothing in the article suggesting they've changed their strategy with TPUs in any way. The word TPU isn't even mentioned here. There's no suggestion they're actually using this internally either. There's no benchmarks showing that it's more cost-effective or scales better. And isn't your second para…

The word TPU isn't even mentioned here. A sentence that reads "I am going to eat nothing but vegetables from now on" doesn't mention meat, but you can infer that I won't eat meat again from the sentence. A sentence that says Google are going all in on nVidea GPUs for AI doesn't need to mention TPUs to convey information about their future either.

That's not true.

Google is huge.

Just a few H100 doesn't represent anything huge in Google scale.

I also tried to find your analogy in that article and google announcement and it's not there.

Post reply on HN