Live data from Hacker News

Ironwood: The first Google TPU for the age of inference

blog.google

11–20 of 186 posts

Re: Ironwood: The first Google TPU for the age of inference

#11

It looks amazing but I wish we could stop playing silly games with benchmarks. Why compare fp8 performance in ironwood to architectures which don't support fp8 in hardware? Why leave out TPUv6 in the comparison? Why compare fp64 flops in the El Capitan supercomputer to fp8 flops in the TPU pod when you know full well these are not comparable? [Edit: it turns out that El Capitan is actually faster when compared like f…

Also, there is no such thing as a "El Capitan pod". The quoted number is for the entire supercomputer.

My impression from this is that they are too scared to say that their TPU pod is equivalent to 60 GB200 NVL72 racks in terms of fp8 flops.

I can only assume that they need way more than 60 racks and they want to hide this fact.

Re: Ironwood: The first Google TPU for the age of inference

#13
Not knowing much about special-purpose chips, I would like to understand whether chips like this would give Google a significant cost advantage over the likes of Anthropic or OpenAI when offering LLM services. Is similar technology available to Google's competitors?

Re: Ironwood: The first Google TPU for the age of inference

#16

It looks amazing but I wish we could stop playing silly games with benchmarks. Why compare fp8 performance in ironwood to architectures which don't support fp8 in hardware? Why leave out TPUv6 in the comparison? Why compare fp64 flops in the El Capitan supercomputer to fp8 flops in the TPU pod when you know full well these are not comparable? [Edit: it turns out that El Capitan is actually faster when compared like f…

>Why compare fp8 performance in ironwood to architectures which don't support fp8 in hardware?

Because end users want to use fp8. Why should architectural differences matter when the speed is what matters at the end of the day?

Re: Ironwood: The first Google TPU for the age of inference

#17
post #4

It looks amazing but I wish we could stop playing silly games with benchmarks. Why compare fp8 performance in ironwood to architectures which don't support fp8 in hardware? Why leave out TPUv6 in the comparison? Why compare fp64 flops in the El Capitan supercomputer to fp8 flops in the TPU pod when you know full well these are not comparable? [Edit: it turns out that El Capitan is actually faster when compared like f…

I think it’s not misleading, but rather very clear that there are problems. v7 is compared to v5e. Also, notice that it’s not compared to competitors, and the price isn’t mentioned. Finally, I think the much bigger issue with TPU is the software and developer experience. Without improvements there, there’s close to zero chance that anyone besides a few companies will use TPU. It’s barely viable if the trend continues…

>Without improvements there, there’s close to zero chance that anyone besides a few companies will use TPU. It’s barely viable if the trend continues.

I wonder whether Google sees this as a problem. In a way it just means more AI compute capacity for Google.

Re: Ironwood: The first Google TPU for the age of inference

#18
post #12

The first specifically designed for inference? Wasn’t the original TPU inference only?

Yup. (Source: was at brain at the time.)

Also holy cow that was 10 years ago already? Dang.

Amusing bit: The first TPU design was based on fully connected networks; the advent of CNNs forced some design rethinking, and then the advent of RNNs (and then transformers) did it yet again.

So maybe it's reasonable to say that this is the first TPU designed for inference in the world where you have both a matrix multiply unit and an embedding processor.

(Also, the first gen was purely a co-processor, whereas the later generations included their own network fabric, a trait shared by this most recent one. So it's not totally crazy to think of the first one as a very different beast.)

Re: Ironwood: The first Google TPU for the age of inference

#19
post #8
post #6

Earlier quoted context omitted.

The reference to El Capitan, is a competitor.

Are you suggesting NVIDIA is not a competitor?

You said: "notice that it’s not compared to competitors"

The article says: "When scaled to 9,216 chips per pod for a total of 42.5 Exaflops, Ironwood supports more than 24x the compute power of the world’s largest supercomputer – El Capitan – which offers just 1.7 Exaflops per pod."

It is literally compared to a competitor.

Post reply on HN