Live data from Hacker News

TPUs vs. GPUs and why Google is positioned to win AI race in the long term

uncoveralpha.com

241–250 of 328 posts

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#241
The part that surprised me is how much TPUs gain from the systolic array design. It basically cuts down the constant memory shuffling that GPUs have to do, so more of the chip’s time is spent actually computing.

The downside is the same thing that makes them fast: they’re very specialized. If your code already fits the TPU stack (JAX/TensorFlow), you get great performance per dollar. If not, the ecosystem gap and fear of lock-in make GPUs the safer default.

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#242
post #215

I always enjoy being wrong and I was very wrong in my predictions about Google : I thought they should theoretically win, but I was also very confident they couldn't possibly turn their execution ship around to actually pull together a coherent competitor to OpenAI. But they do seem to have done that and it's very impressive. If they do continue to execute, I can't see anybody stopping them dominating and I would be…

> because of the lack of any clear or reasonable statement or guidelines on how they use your data.

They’ve been very clear, in my opinion: https://cloud.google.com/gemini/docs/discover/data-governanc...

I suppose there will always be the people who refuse to trust them or choose to believe they’re secretly doing something different.

However I’m not sure what you’re referring to by saying they haven’t said anything about how data is used.

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#243

All this assumes that LLMs are the sole mechanism for AI and will remain so forever: no novel architectures (neither hardware nor software), no progress in AI theory, nothing better than LLMs, simply brute force LLM computation ad infinitum . Perhaps the assumptions are true. The mere presence of LLMs seems to have lowered the IQ of the Internet drastically, sopping up financial investors and resources that might oth…

TPUs predate LLMs by a long time. They were already being used for all the other internal ML work needed for search, youtube, etc.

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#244

Earlier quoted context omitted.

Nothing in principle. But Huang probably doesn't believe in hyper specializing their chips at this stage because it's unlikely that the compute demands of 2035 are something we can predict today. For a counterpoint, Jim Keller took Tenstorrent in the opposite direction. Their chips are also very efficient, but even more general purpose than NVIDIA chips.

How is Tenstorrent h/w more general purpose than NVIDIA chips? TT hardware is only good for matmuls and some elementwise operations, and plain sucks for anything else. Their software is abysmal.

Of course there's the general purpose RISC V CPU controller component but also, each NPU is designed in troikas that have one core reading data in, one core performing the actual kernel work, and the third core forwarding data out.

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#245

I don't think what the article writes about matters all that much. Gemini 3 Pro is arguably not even the best model anymore, and it's _weeks_ old, and Google has far more resources than Anthropic does. If the hardware actually was the secret sauce, Google would be wiping the floor with little everyone else. But they're not. There's a few confounding problems: 1. Actually using that hardware effectively isn't easy. It…

_Weeks_ old! What a fossil! Slightly more seriously: what you say makes sense if and only if you're projecting Sam Altman and assuming that a) real legit superhuman AGI is just around the corner, and b) all the spoils will accrue to the first company that finds it, which means you need to be 100% in on building the next model that will finally unlock AGI. But if this is not the case -- and it's increasingly looking l…

> _Weeks_ old! What a fossil!

I think you are missing the point. They are saying "weeks old" isn't very old.

> it's going to continue to be a race of competing AIs, and that race will be won by the company that can deliver AI at scale the most cheaply.

I don't see how that follows at all. Quality and distribution both matter a lot here.

Google has some advantages but some disadvantages here too.

If you are on AWS GovCloud, Anthropic is right there. Same on Azure, and on Oracle.

I believe Gemini will be available on the Oracle Cloud at some point (it has been announced) but they are still behind in the enterprise distribution race.

OpenAI is only available on Azure, although I believe their new contract lets them strike deals elsewhere.

On the consumer side, OpenAI and Google are well ahead of course.

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#246
post #218

If Google won, it would cannibalize its current ad-driven business and replace it with something that is extremely expensive to run and difficult to make profit from. A Pyrrhic win essentially.

I mean, focus is a thing that Google has always struggled with. But I kind of doubt that customers who need online marketing (ads) are going to convert en masse to users who rent cloud TPUs instead.

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#247
post #239

Earlier quoted context omitted.

> Almost all players are running Nvidia other than Google. No surprises there, Google is not the greatest company at productizing their tech for external consumption. > The other players are certainly more than just competing with Google. TBF, its easy to stay in the game when you're flush with cash, and for the past N-quarters, investors have been throwing money at AI companies, Nvidia's margins have greatly benefit…

Ok? The person I was replying to was saying that Google’s compute offering is substantially superior to Nvidia’s. What do your comments about market positioning have to do with that? If Google’s TPUs were really substantially superior, don’t you think that would result in at least short term market advantages for Gemini? Where are they?

They are suggesting it is easier for others to buy buy more NVidia chips and feed them more power. Whilst operating costs are covered by investors. If they move on to competing on having to do inference the cheepest then the TPUs will shine.

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#248
post #41

Google's real moat isn't the TPU silicon itself—it's not about cooling, individual performance, or hyper-specialization—but rather the massive parallel scale enabled by their OCS interconnects. To quote The Next Platform: "An Ironwood cluster linked with Google’s absolutely unique optical circuit switch interconnect can bring to bear 9,216 Ironwood TPUs with a combined 1.77 PB of HBM memory... This makes a rackscale…

That is comparing an all to all switched Nvlink fabric to a 3D torus for TPUs. Those are completely different network topologies with different tradeoffs. For example the currently very popular Mixture of Experts architectures require a lot of all to all traffic (for expert parallelism) which works a lot better on the switched NVlink fabric as opposed where it doesn't need to traverse multiple links in the torus.

Really? Fully-connected hardware is in buildable (at scale) which we already know from the HPC world. Fat trees and dragonfly networks are pretty scalable, but a 3d torus is a very good tradeofff, and respects the dimensionality of reality.

Bisection bandwidth is a useful metric, but is hop count? Per-hop cost tends to be pretty small.

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#249
post #177
post #89

Earlier quoted context omitted.

Check the specs again. Per chip, TPU 7x has 192GB of HBM3e, whereas the NVIDIA B200 has 186GB. While the B200 wins on raw FP8 throughput (~9000 vs 4614 TFLOPs), that makes sense given NVIDIA has optimized for the single-chip game for over 20 years. But the bottleneck here isn't the chip—it's the domain size. NVIDIA's top-tier NVL72 tops out at an NVLink domain of 72 Blackwell GPUs. Meanwhile, Google is connecting 921…

Wow, no, not at all. It’s better to have a set of smaller, faster cliques connected by a slow network than a slower-than-clique flat network that connects everything. The cliques connected by a slow DCN can scale to arbitrary size. Even Google has had to resort to that for its biggest clusters.

Is this claim based on observed comm patterns in some particular AI architecture?

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#250

I don't think what the article writes about matters all that much. Gemini 3 Pro is arguably not even the best model anymore, and it's _weeks_ old, and Google has far more resources than Anthropic does. If the hardware actually was the secret sauce, Google would be wiping the floor with little everyone else. But they're not. There's a few confounding problems: 1. Actually using that hardware effectively isn't easy. It…

_Weeks_ old! What a fossil! Slightly more seriously: what you say makes sense if and only if you're projecting Sam Altman and assuming that a) real legit superhuman AGI is just around the corner, and b) all the spoils will accrue to the first company that finds it, which means you need to be 100% in on building the next model that will finally unlock AGI. But if this is not the case -- and it's increasingly looking l…

> _Weeks_ old! What a fossil!

Last week it looked like Google had won (hence the blog post) but now almost nobody is talking about antigravity and Gemini 3 anymore so yeah what op says is relevant

Post reply on HN