Live data from Hacker News

Ironwood: The first Google TPU for the age of inference

blog.google

151–160 of 186 posts

Re: Ironwood: The first Google TPU for the age of inference

#153
"Age of inference", "age of Gemini", bleh. Google's doing some amazing work but I hate how they communicate - vague, self-serving, low signal-to-noise slop.

It's like they take some interesting wood carving of communication, then sand it down to a featureless nub.

Re: Ironwood: The first Google TPU for the age of inference

#154
post #75

Earlier quoted context omitted.

Nvidia has ~60% margins in their datacenter chips. So TPU's have quite a bit of headroom to save google money without being as good as Nvidia GPU's. No one else has access to anything similar, Amazon is just starting to scale their Trainium chip.

Microsoft has the MAIA 100 as well. No comment on their scale/plans though.

[deleted]

Re: Ironwood: The first Google TPU for the age of inference

#157
post #110
post #102

Earlier quoted context omitted.

Wouldn't a 3d torus network have horrible performance with 9,216 nodes? And really horrible latency? I'd have assumed traditional spine-leaf would do better. But I must be wrong as they're claiming their latency is great here. Of course, they provide zero actual evidence of that. And I'll echo, what even is an AI data center, because we're still none the wiser.

> what even is an AI data center A data center that runs significant AI training or inference loads. Non AI data centers are fairly commodity. Google's non-AI efficiency is not much better than Amazon or anyone else. Google is much more efficient at running AI workloads than anyone else.

> Google's non-AI efficiency is not much better than Amazon or anyone else.

I don't think this is true. Google has long been a leader in efficiency. Look at the power usage effectiveness (PUE). A decade ago Google announced average PUEs around 1.12 while the industry average was closer to 2.0. From what I can tell they reported a 1.1 average fleet wide last year. They've been more transparent about this than any of the other big players.

AWS is opaque by comparison, but they report 1.2 on average. So they're close now, but that's after a decade of trying to catch up to Google.

To suggest the rest of the industry is on the same level is not at all accurate.

https://en.wikipedia.org/wiki/Power_usage_effectiveness

(Amazon isn't even listed in the "Notably efficient companies" section on the Wikipedia page).

Re: Ironwood: The first Google TPU for the age of inference

#158
post #12

The first specifically designed for inference? Wasn’t the original TPU inference only?

The phrasing is very precise here, it’s the first TPU for _the age of inference_, which is a novel marketing term they have defined to refer to CoT and Deep Research.

Ugh. We should have caught that.

Can anyone suggest a better (i.e. more accurate and neutral) title, devoid of marketing tropes?

Re: Ironwood: The first Google TPU for the age of inference

#159
post #12

The first specifically designed for inference? Wasn’t the original TPU inference only?

The first one was designed as a proof of concept that it would work at all, not really to be optimal for inference workloads. It just turns out that inference is easier.
Post reply on HN