Live data from Hacker News

TPUs vs. GPUs and why Google is positioned to win AI race in the long term

uncoveralpha.com

231–240 of 328 posts

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#231
post #45
post #10

Earlier quoted context omitted.

That's exactly what Nvidia is doing with tensor cores.

Except the native width of Tensor Cores are about 8-32 (depending on scalar type), whereas the width of TPUs is up to 256. The difference in scale is massive.

If it turns out to be useful, Nvidia can't just tweak a parameter in their verilog and declare victory?

If not, what's fundamentally difficult about doing 32 vs 256 here?

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#232

In my 20+ years of following NVIDIA, I have learned to never bet against them long-term. I actually do not know exactly why they continually win, but they do. The main issue they have a 3-4 year gap between wanting a new design pivot and realizing it (silicon has a long "pipeline"), it can seem that they may be missing a new trend or swerve in the demands of the market, it is often simply because there is this delay.

Turkeys bet on tomorrow 364 days of the year.

I told you a thousand times, you have to sell your pumpkin stock before Halloween, before!

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#233

I don't think what the article writes about matters all that much. Gemini 3 Pro is arguably not even the best model anymore, and it's _weeks_ old, and Google has far more resources than Anthropic does. If the hardware actually was the secret sauce, Google would be wiping the floor with little everyone else. But they're not. There's a few confounding problems: 1. Actually using that hardware effectively isn't easy. It…

They are using that hardware to wipe the floor with everyone if you look at the price per million tokens.

Gemini3 is slightly more expensive than GPT5.1 for both input and output tokens though?

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#234
post #190
post #89

Earlier quoted context omitted.

Check the specs again. Per chip, TPU 7x has 192GB of HBM3e, whereas the NVIDIA B200 has 186GB. While the B200 wins on raw FP8 throughput (~9000 vs 4614 TFLOPs), that makes sense given NVIDIA has optimized for the single-chip game for over 20 years. But the bottleneck here isn't the chip—it's the domain size. NVIDIA's top-tier NVL72 tops out at an NVLink domain of 72 Blackwell GPUs. Meanwhile, Google is connecting 921…

I guess “this weight class” is some theoretical class divorced from any application? Almost all players are running Nvidia other than Google. The other players are certainly more than just competing with Google.

> Almost all players are running Nvidia other than Google.

No surprises there, Google is not the greatest company at productizing their tech for external consumption.

> The other players are certainly more than just competing with Google.

TBF, its easy to stay in the game when you're flush with cash, and for the past N-quarters, investors have been throwing money at AI companies, Nvidia's margins have greatly benefited from this largesse. There will be blood on the floor once investors start demanding returns to their investments.

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#235
post #41

Google's real moat isn't the TPU silicon itself—it's not about cooling, individual performance, or hyper-specialization—but rather the massive parallel scale enabled by their OCS interconnects. To quote The Next Platform: "An Ironwood cluster linked with Google’s absolutely unique optical circuit switch interconnect can bring to bear 9,216 Ironwood TPUs with a combined 1.77 PB of HBM memory... This makes a rackscale…

NVFP4 is the thing no one saw coming. I wasn't watching the MX process really, so I cast no judgements, but it's exactly what it sounds like, a serious compromise in resource constrained settings. And it's in the silicon pipeline.

NVFP4 is to put it mildly a masterpiece, the UTF-8 of its domain and in strikingly similar ways it is 1. general 2. robust to gross misuse 3. not optional if success and cost both matter.

It's not a gap that can be closed by a process node or an architecture tweak: it's an order of magnitude where the polynomials that were killing you on the way up are now working for you.

sm_120 (what NVIDIA's quiet repos call CTA1) consumer gear does softmax attention and projection/MLP blockscaled GEMM at a bit over a petaflop at 300W and close to two (dense) at 600W.

This changes the whole game and it's not clear anyone outside the lab even knows the new equilibrium points, it's nothing like Flash3 on Hopper, lotta stuff looks FLOPs bound, GDDR7 looks like a better deal than HBMe3. The DGX Spark is in no way deficient, it has ample memory bandwidth.

This has been in the pipe for something like five years and even if everyone else started at the beginning of the year when this was knowable, it would still be 12-18 months until tape out. And they haven't started.

Years Until Anyone Can Compete With NVIDIA is back up to the 2-5 it was 2-5 years ago.

This was supposed to be the year ROCm and the new Intel stuff became viable.

They had a plan.

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#236
post #28
post #11

That and the fact they can self-fund the whole AI venture and don't require outside investment.

That and they were harvesting data way before it was cool, and now that it is cool, they're in a privileged position since almost no-one can afford to block GoogleBot. They do voluntarily offer a way to signal that the data GoogleBot sees is not to be used for training, for now, and assuming you take them at their word, but AFAIK there is no way to stop them doing RAG on your content without destroying your SEO in th…

But they also collect the data without causing denial of service, and respect robots.txt, which is more than you can say of most LLM scrapers...

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#237

Earlier quoted context omitted.

Isn’t there a suspicion that OpenAI buying custom chips from another Sam Altman venture is just graft? Wasn’t that one of the things that came up when the board tried to out him?

The chips are being done in-house.

It was only brought in-house after the $5,000,000,000,000 self-dealing AI chip venture failed to launch.

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#238
deepseek kind of innovated on this using off-the-shelf components right ?

to quote from their paper "In order to ensure sufficient computational performance for DualPipe, we customize efficient cross-node all-to-all communication kernels (including dispatching and combining) to conserve the number of SMs dedicated to communication. The implementation of the kernels is codesigned with the MoE gating algorithm and the network topology of our cluster."

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#239
post #190

Earlier quoted context omitted.

I guess “this weight class” is some theoretical class divorced from any application? Almost all players are running Nvidia other than Google. The other players are certainly more than just competing with Google.

> Almost all players are running Nvidia other than Google. No surprises there, Google is not the greatest company at productizing their tech for external consumption. > The other players are certainly more than just competing with Google. TBF, its easy to stay in the game when you're flush with cash, and for the past N-quarters, investors have been throwing money at AI companies, Nvidia's margins have greatly benefit…

Ok? The person I was replying to was saying that Google’s compute offering is substantially superior to Nvidia’s. What do your comments about market positioning have to do with that?

If Google’s TPUs were really substantially superior, don’t you think that would result in at least short term market advantages for Gemini? Where are they?

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#240
post #20

Earlier quoted context omitted.

They will, I'm sure. The big difference is that Google is both the chip designer *and* the AI company. So they get both sets of profits. Both Google and Nvidia contract TSMC for chips. Then Nvidia sells them at a huge profit. Then OpenAI (for example) buys them at that inflated rate and them puts them into production. So while Nvidia is "selling shovels", Google is making their own shovels and has their own mines.

So when the bubble pops the companies making the shovels (TSMC, NVIDIA) might still have the money they got for their products and some of the ex-AI companies might least be able to sell standard compliant GPUs on the wider market. And Google will end up with lots of useless super specialized custom hardware.

Google uses TPUs for its internal AI work (training Gemini for example), which surely isn't decreasing in demand or usage as their portfolio and product footprint increases. So I have a feeling they'd be able to put those TPUs to good use?
Post reply on HN