Live data from Hacker News

TPUs vs. GPUs and why Google is positioned to win AI race in the long term

uncoveralpha.com

281–290 of 328 posts

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#281

Google has always had great tech - their problem is the product or the perseverance, conviction, and taste needed to make things people want.

Their incentive structure doesn't lead to longevity. Nobody gets promoted for keeping a product alive, they get promoted for shipping something new. That's why we're on version 37 of whatever their chat client is called now. I think we can be reasonably sure that search, Gmail, and some flavor of AI will live on, but other than that, Google apps are basically end-of-life at launch.

Google released their latest chat app 8 years ago.

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#282
post #20
post #2

A question I don't see addressed in all these articles: what prevents Nvidia from doing the same thing and iterating on their more general-purpose GPU towards a more focused TPU-like chip as well, if that turns out to be what the market really wants.

They will, I'm sure. The big difference is that Google is both the chip designer *and* the AI company. So they get both sets of profits. Both Google and Nvidia contract TSMC for chips. Then Nvidia sells them at a huge profit. Then OpenAI (for example) buys them at that inflated rate and them puts them into production. So while Nvidia is "selling shovels", Google is making their own shovels and has their own mines.

The own shovels for own mines strategy has a hidden downside: isolation. NVIDIA sells shovels to everyone - OpenAI, Meta, xAI, Microsoft - and gets feedback from the entire market. They see where the industry is heading faster than Google, which is stewing in its own juices. While Google optimizes TPUs for current Google tasks (Gemini, Search), NVIDIA optimizes GPUs for all possible future tasks. In an era of rapid change, the market's hive mind usually beats closed vertical integration.

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#283
post #41

Google's real moat isn't the TPU silicon itself—it's not about cooling, individual performance, or hyper-specialization—but rather the massive parallel scale enabled by their OCS interconnects. To quote The Next Platform: "An Ironwood cluster linked with Google’s absolutely unique optical circuit switch interconnect can bring to bear 9,216 Ironwood TPUs with a combined 1.77 PB of HBM memory... This makes a rackscale…

NVFP4 is the thing no one saw coming. I wasn't watching the MX process really, so I cast no judgements, but it's exactly what it sounds like, a serious compromise in resource constrained settings. And it's in the silicon pipeline. NVFP4 is to put it mildly a masterpiece, the UTF-8 of its domain and in strikingly similar ways it is 1. general 2. robust to gross misuse 3. not optional if success and cost both matter. I…

This comment reads as if it were LLM-generated.

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#284

Earlier quoted context omitted.

NVFP4 is the thing no one saw coming. I wasn't watching the MX process really, so I cast no judgements, but it's exactly what it sounds like, a serious compromise in resource constrained settings. And it's in the silicon pipeline. NVFP4 is to put it mildly a masterpiece, the UTF-8 of its domain and in strikingly similar ways it is 1. general 2. robust to gross misuse 3. not optional if success and cost both matter. I…

This comment reads as if it were LLM-generated.

Agree.

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#285

I don't think what the article writes about matters all that much. Gemini 3 Pro is arguably not even the best model anymore, and it's _weeks_ old, and Google has far more resources than Anthropic does. If the hardware actually was the secret sauce, Google would be wiping the floor with little everyone else. But they're not. There's a few confounding problems: 1. Actually using that hardware effectively isn't easy. It…

Google owns 14% of Anthropic and Anthropic is using Google TPUs, as well as AWS Trainium and of course GPUs. It isn't necessary for one company to create both the winning hardware and the winning software to be part of the solution. In fact with the close race in software hardware seems like the better bet.

https://www.anthropic.com/news/expanding-our-use-of-google-c...

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#286

Earlier quoted context omitted.

That is comparing an all to all switched Nvlink fabric to a 3D torus for TPUs. Those are completely different network topologies with different tradeoffs. For example the currently very popular Mixture of Experts architectures require a lot of all to all traffic (for expert parallelism) which works a lot better on the switched NVlink fabric as opposed where it doesn't need to traverse multiple links in the torus.

Really? Fully-connected hardware is in buildable (at scale) which we already know from the HPC world. Fat trees and dragonfly networks are pretty scalable, but a 3d torus is a very good tradeofff, and respects the dimensionality of reality. Bisection bandwidth is a useful metric, but is hop count? Per-hop cost tends to be pretty small.

Latency (of different types), jitter, and guaranteed bandwidth are the real underlying metrics. Hop count is just one potential driver of those, but different approaches may or may not tackle each of these parts differently.

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#287
post #80

Earlier quoted context omitted.

> I'd rather have a winning google than openai or meta anyway. Why? To me, it seems better for the market, if the best models and the best hardware were not controlled by the same company.

I agree, it would be the best of bad cases, in a sense. I have low trust in OpenAI due to its leadership, and in Meta, because, well, Meta has history, let's say.

I think you are disagreeing.

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#288
post #80

Earlier quoted context omitted.

I agree, it would be the best of bad cases, in a sense. I have low trust in OpenAI due to its leadership, and in Meta, because, well, Meta has history, let's say.

I think you are disagreeing.

Yep :)

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#289
post #79

Earlier quoted context omitted.

The best path for victory for China is probably no war at all. War is wasteful and risky.

In a world where those starting wars would suffer their consequences the most, wars would be a bad idea. This is not such a world.

I don't agree.

Even the US suffers (their veterans do, anyway) but that's the country that in general least suffers from their constant involvement in warfare. They have this industry down to a "T". However, you cannot generalize from a nation that can project military force almost anywhere in the globe, with little fear of repercussion back home; most countries cannot afford this. China certainly cannot.

So what about the rest? Internecine conflicts are outrageously wasteful, and sadly common in the modern age. Russia's war with Ukraine has turned incredibly wasteful and costly, and Russians are suffering (and dying) regardless of whatever Putin says.

I think China is not generally oriented towards waging war. They do have their military, military projects, and their nationalistic things (what I learn from Wikipedia is called "irredentism"), but generally they seem to be trying to become an economic world power. War would mess and interfere with that. War is too fucking risky.

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#290

Earlier quoted context omitted.

It's not that the TPU is better than an NVidia GPU, it's just that it's cheaper since it doesn't have a fat NVidia markup applied, and is also better vertically integrated since it was designed/specified by Google for Google.

TPUs are also cheaper because GPUs need to be more general purpose whereas TPUs are designed with a focus on LLM workloads meaning there's not wasted silicon. Nothing's there that doesn't need to be there. The potential downside would be if a significantly different architecture arises that would be difficult for TPUs to handle and easier for GPUs (given their more general purpose). But even then Google could probabl…

The T in TPU stands for tensor, which in this context is just a fancy matrix. These days both are optimised for matrix algebra, i.e. general ML workloads, not just LLMs.

If LLMs become unfashionable they’ll still be good for other ML tasks like image recognition.

Post reply on HN