Live data from Hacker News

Ironwood: The first Google TPU for the age of inference

blog.google

141–150 of 186 posts

Re: Ironwood: The first Google TPU for the age of inference

#143
post #132

Earlier quoted context omitted.

Cerebras (and Groq) has the problem of using too much die for compute and not enough for memory. Their method of scaling is to fan out the compute across more physical space. This takes more dc space, power and cooling, which is a huge issue. Funny enough, when I talked to Cerebras at SC24, they told me their largest customers are for training, not inference. They just market it as an inference product, which is even…

Thank you for sharing this perspective — really insightful. I’ve been reading up on Groq’s architecture and was under the impression that their chips dedicate a significant portion of die area to on-chip SRAM (around 220MiB per chip, if I recall correctly), which struck me as quite generous compared to typical accelerators. From die shots and materials I’ve seen, it even looks like ~40% of the die might be allocated…

> is it a matter of absolute capacity still being insufficient for current model sizes

This. Additionally, models aren't getting smaller, they are getting bigger and to be useful to a wider range of users, they also need more context to go off of, which is even more memory.

Previously: https://news.ycombinator.com/item?id=42003823

It could be partially the DC, but look at the rack density... to get to an equal amount of GPU compute and memory, you need 10x the rack space...

https://www.linkedin.com/posts/andrewdfeldman_a-few-weeks-ag...

Previously: https://news.ycombinator.com/item?id=39966620

Now compare that to an NV72 and the direction Dell/CoreWeave/Switch are going in with the EVO containment... far better. One can imagine that AMD might do something similar.

https://www.coreweave.com/blog/coreweave-pushes-boundaries-w...

Re: Ironwood: The first Google TPU for the age of inference

#144
post #131

I wonder if these chips might contribute towards advancements for the Coral TPU chips?

There were pretty much abandoned years ago. The software stack support from Google only really lasted a few months and even then the stack they ran on was already years old versions of operating systems and python versions.

The only support is via a few enthusiastic third party developers.

Re: Ironwood: The first Google TPU for the age of inference

#145
post #73

It looks amazing but I wish we could stop playing silly games with benchmarks. Why compare fp8 performance in ironwood to architectures which don't support fp8 in hardware? Why leave out TPUv6 in the comparison? Why compare fp64 flops in the El Capitan supercomputer to fp8 flops in the TPU pod when you know full well these are not comparable? [Edit: it turns out that El Capitan is actually faster when compared like f…

Google shouldn't do that comparison. When I worked there I strongly emphasized to the TPU leadership to not compare their systems to supercomputers- not only were the comparisons misleading, Google absolutely does not want supercomputer users to switch to TPUs. SC users are demanding and require huge support.

Google needs to sell to Enterprise Customers. It's a Google Cloud Event. Of course they have incentives to hype because once long-term contracts are signed you lose that customer forever. So, hype is a necessity

Re: Ironwood: The first Google TPU for the age of inference

#146

Earlier quoted context omitted.

I like owning things

You will own nothing and you will be happy.

Seriously yes. I don’t want to own rapidly depreciating hardware if I don’t have to.

I don’t want to own open source software and I’d prefer if culture was derived from and contributed to the public domain.

I’m still fully on board with capitalism, but there are many instances where I’d prefer to replace physical acquisition with renting, or replacing corporate-made culture with public culture.

Re: Ironwood: The first Google TPU for the age of inference

#147
post #130

Earlier quoted context omitted.

> Nvidia seemed 'untouchable' for so long in this space that its nice to see things get shaken up. Did it? Both Mistral's LeChat (running on Cerebras) and Google's Gemini (running on Tensors) have clearly showed ages ago Nvidia had no advantage at all in inference. The hundreds of billions spent in hardware till now focused on training, but inference is in the long run gonna get the lion share of the work.

> but inference is in the long run gonna get the lion share of the work. I'm not sure - might not the equilibrium state be that we are constantly fine-tuning models with the latest data (e.g. social media firehose)?

Head of groq said that in his experience at google training was less than 10% of compute.

Re: Ironwood: The first Google TPU for the age of inference

#148

Earlier quoted context omitted.

You can't get excited about lower prices for your cloud GPU workloads thanks to the competition it brings to Nvidia? This benefits everyone, even if you don't use Google Cloud, because of the competition it introduces.

I like owning things

[deleted]

Re: Ironwood: The first Google TPU for the age of inference

#150

Earlier quoted context omitted.

You will own nothing and you will be happy.

Seriously yes. I don’t want to own rapidly depreciating hardware if I don’t have to. I don’t want to own open source software and I’d prefer if culture was derived from and contributed to the public domain. I’m still fully on board with capitalism, but there are many instances where I’d prefer to replace physical acquisition with renting, or replacing corporate-made culture with public culture.

Nvidia cloud instances have competition. What happens when you get vendor lock in with TPU and have no exit plan? Arguably it's competition that drives value creation, not capitalism.
Post reply on HN