Live data from Hacker News

TPUs vs. GPUs and why Google is positioned to win AI race in the long term

uncoveralpha.com

251–260 of 328 posts

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#251

> It is also important to note that, until recently, the GenAI industry’s focus has largely been on training workloads. In training workloads, CUDA is very important, but when it comes to inference, even reasoning inference, CUDA is not that important, so the chances of expanding the TPU footprint in inference are much higher than those in training (although TPUs do really well in training as well – Gemini 3 the prim…

This is a very important point - the market for training chips might be a bubble, but the market for inference is much, much larger. At some point we might have good enough models and the need for new frontier models will cool down. The big power-hungry datacenters we are seeing are mostly geared towards training, while inference-only systems are much simpler and power efficient. A real shame, BTW, all that silicon d…

Some more traditional number crunching has long looked at lower- and mixed-precision hardware.

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#252
post #247
post #239

Earlier quoted context omitted.

Ok? The person I was replying to was saying that Google’s compute offering is substantially superior to Nvidia’s. What do your comments about market positioning have to do with that? If Google’s TPUs were really substantially superior, don’t you think that would result in at least short term market advantages for Gemini? Where are they?

They are suggesting it is easier for others to buy buy more NVidia chips and feed them more power. Whilst operating costs are covered by investors. If they move on to competing on having to do inference the cheepest then the TPUs will shine.

The original post made no comments about inference or training or even cost in any way. It said you could hook up more TPUs together with more memory and higher average bandwidth than you could with a datacenter of Nvidia GPUs. From an architectural point of view, it isn’t clear (and is not explained) what that enables. It clearly hasn’t led to a business outcome for Google where they are the clear market leader.

Seemingly fast interconnects benefit training more than inference since training can have more parallel communication between nodes. Inference for users is more embarrassingly parallel (requires less communication) than updating and merging network weights.

My point: cool benchmark, what does it matter? The original post says Nvidia doesn’t have anything to compete with massively interconnected TPUs. It didn’t merely say Google’s TPUs were better. It said that Nvidia can’t compete. That’s clearly bullshit and wishful thinking, right? There is no evidence in the market to support that, and no actual technical points have been presented in this thread either. OpenAI, Anthropic, etc are certainly competing with Google, right?

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#253
post #41

Google's real moat isn't the TPU silicon itself—it's not about cooling, individual performance, or hyper-specialization—but rather the massive parallel scale enabled by their OCS interconnects. To quote The Next Platform: "An Ironwood cluster linked with Google’s absolutely unique optical circuit switch interconnect can bring to bear 9,216 Ironwood TPUs with a combined 1.77 PB of HBM memory... This makes a rackscale…

For all the excitement surrounding this, I fail to comprehend how Google can't even meet the current demand for Gemini 3^. Moreover, they are unwilling to invest in expansion directly (apparently have a mandate to double their compute every 6 months without spending more than their current budget). So, pardon me if I can't see how they will scale operations as demand grows while simultaneously selling their chips to competitors?! This situation doesn't make any sense.

^Even now I get capacity related error messages, so many days after the Gemini 3 launch. Also, Jules is basically unusable. Maybe Gemini 3 is a bigger resource hog than anyone outside of Google realizes.

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#254
post #41

Google's real moat isn't the TPU silicon itself—it's not about cooling, individual performance, or hyper-specialization—but rather the massive parallel scale enabled by their OCS interconnects. To quote The Next Platform: "An Ironwood cluster linked with Google’s absolutely unique optical circuit switch interconnect can bring to bear 9,216 Ironwood TPUs with a combined 1.77 PB of HBM memory... This makes a rackscale…

For all the excitement surrounding this, I fail to comprehend how Google can't even meet the current demand for Gemini 3^. Moreover, they are unwilling to invest in expansion directly (apparently have a mandate to double their compute every 6 months without spending more than their current budget). So, pardon me if I can't see how they will scale operations as demand grows while simultaneously selling their chips to…

I also suspect Google is launching models it can’t really sustain in volume or that are operating at a loss. Nothing preventing them from like doubling model size compared to the rest or allocating an insane amount of compute just to make the headlines on model performance (clearly it’s good for the stock). These things are opaque anyway, buried deep into the P&L.

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#255
post #218

If Google won, it would cannibalize its current ad-driven business and replace it with something that is extremely expensive to run and difficult to make profit from. A Pyrrhic win essentially.

But that would happen regardless of who won, better to at least dominate the new paradigm and figure out how to extract value from it. I also suspect that once the value generation is figured out they will cease offering these APIs to anyone, if you had a golden goose would you rent it?

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#256
post #215

I always enjoy being wrong and I was very wrong in my predictions about Google : I thought they should theoretically win, but I was also very confident they couldn't possibly turn their execution ship around to actually pull together a coherent competitor to OpenAI. But they do seem to have done that and it's very impressive. If they do continue to execute, I can't see anybody stopping them dominating and I would be…

The LLM provider I trust the most right now is AWS. Anybody else seems to have very conflicted purposes when it comes to sending them my data and interactions.

[deleted]

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#257

Earlier quoted context omitted.

Their incentive structure doesn't lead to longevity. Nobody gets promoted for keeping a product alive, they get promoted for shipping something new. That's why we're on version 37 of whatever their chat client is called now. I think we can be reasonably sure that search, Gmail, and some flavor of AI will live on, but other than that, Google apps are basically end-of-life at launch.

It's also paradoxically the talent in tech that isolates them. The internal tech stack is so incredibly specialized, most Google products have to either be built for internal users or external users. Agree there are lots of other contributing causes like culture, incentives, security, etc.

haha remember Steve Yegge's Platform Rant? [1]

nothing changed...

[1] Ref: he mistakenly posted what was meant to be an internal memo, publicly on G+. He quickly took it down but of course The Internet Never Forgets https://gist.github.com/chitchcock/1281611

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#258
post #79

Earlier quoted context omitted.

First, the US has advanced fab capabilities and in case of a need can develop them further. On the other side, China will suffer a Russia style blockback while caught up in a nasty war with Taiwan. Totally possible, but the second order effects are much more complex than "leader once for all". The path for victory for China is not war despite the west, but a war when the west would not care.

The best path for victory for China is probably no war at all. War is wasteful and risky.

In a world where those starting wars would suffer their consequences the most, wars would be a bad idea. This is not such a world.

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#259

I don't think what the article writes about matters all that much. Gemini 3 Pro is arguably not even the best model anymore, and it's _weeks_ old, and Google has far more resources than Anthropic does. If the hardware actually was the secret sauce, Google would be wiping the floor with little everyone else. But they're not. There's a few confounding problems: 1. Actually using that hardware effectively isn't easy. It…

On point 5, I think this is the real moat for CUDA. Does Google have tools to optimize kernels on their TPUs? Do they have tools to optimize successive kernel launches on their TPUs? How easy is it to debug on a TPU(arguably CUDA could use work here but still...)? Does Google help me fully utilize their TPUs? Can I warm up a model on a TPU, checkpoint it, and launch the checkpoints to save time?

I am fairly pro-google(they invented the LLM, FFS...) and recognize the advantages(price/token, efficiency, vertical integration, established DCs w/ power allocations) but also know they have a habit of slightly sucking at everything but search.

Re: TPUs vs. GPUs and why Google is positioned to win AI race in the long term

#260
post #82

This is highly relevant: "Meta in talks to spend billions on Google's chips, The Information reports" https://www.reuters.com/business/meta-talks-spend-billions-g...

Weird they'd do this after developing several generations of their own inference chip. Google is basically a competitor. This may just be a ploy to get better pricing from Nvidia.
Post reply on HN