Who manufactures their TPUs? Wondering whether they're US-made, and therefore immune to the tariff craziness.
Ironwood: The first Google TPU for the age of inference
141–150 of 186 posts
Re: Ironwood: The first Google TPU for the age of inference
#142Who manufactures their TPUs? Wondering whether they're US-made, and therefore immune to the tariff craziness.
Re: Ironwood: The first Google TPU for the age of inference
#143Earlier quoted context omitted.
Cerebras (and Groq) has the problem of using too much die for compute and not enough for memory. Their method of scaling is to fan out the compute across more physical space. This takes more dc space, power and cooling, which is a huge issue. Funny enough, when I talked to Cerebras at SC24, they told me their largest customers are for training, not inference. They just market it as an inference product, which is even…
Thank you for sharing this perspective — really insightful. I’ve been reading up on Groq’s architecture and was under the impression that their chips dedicate a significant portion of die area to on-chip SRAM (around 220MiB per chip, if I recall correctly), which struck me as quite generous compared to typical accelerators. From die shots and materials I’ve seen, it even looks like ~40% of the die might be allocated…
This. Additionally, models aren't getting smaller, they are getting bigger and to be useful to a wider range of users, they also need more context to go off of, which is even more memory.
Previously: https://news.ycombinator.com/item?id=42003823
It could be partially the DC, but look at the rack density... to get to an equal amount of GPU compute and memory, you need 10x the rack space...
https://www.linkedin.com/posts/andrewdfeldman_a-few-weeks-ag...
Previously: https://news.ycombinator.com/item?id=39966620
Now compare that to an NV72 and the direction Dell/CoreWeave/Switch are going in with the EVO containment... far better. One can imagine that AMD might do something similar.
https://www.coreweave.com/blog/coreweave-pushes-boundaries-w...
Re: Ironwood: The first Google TPU for the age of inference
#144I wonder if these chips might contribute towards advancements for the Coral TPU chips?
The only support is via a few enthusiastic third party developers.
Re: Ironwood: The first Google TPU for the age of inference
#145It looks amazing but I wish we could stop playing silly games with benchmarks. Why compare fp8 performance in ironwood to architectures which don't support fp8 in hardware? Why leave out TPUv6 in the comparison? Why compare fp64 flops in the El Capitan supercomputer to fp8 flops in the TPU pod when you know full well these are not comparable? [Edit: it turns out that El Capitan is actually faster when compared like f…
Google shouldn't do that comparison. When I worked there I strongly emphasized to the TPU leadership to not compare their systems to supercomputers- not only were the comparisons misleading, Google absolutely does not want supercomputer users to switch to TPUs. SC users are demanding and require huge support.
Re: Ironwood: The first Google TPU for the age of inference
#146Earlier quoted context omitted.
I like owning things
You will own nothing and you will be happy.
I don’t want to own open source software and I’d prefer if culture was derived from and contributed to the public domain.
I’m still fully on board with capitalism, but there are many instances where I’d prefer to replace physical acquisition with renting, or replacing corporate-made culture with public culture.
Re: Ironwood: The first Google TPU for the age of inference
#147Earlier quoted context omitted.
> Nvidia seemed 'untouchable' for so long in this space that its nice to see things get shaken up. Did it? Both Mistral's LeChat (running on Cerebras) and Google's Gemini (running on Tensors) have clearly showed ages ago Nvidia had no advantage at all in inference. The hundreds of billions spent in hardware till now focused on training, but inference is in the long run gonna get the lion share of the work.
> but inference is in the long run gonna get the lion share of the work. I'm not sure - might not the equilibrium state be that we are constantly fine-tuning models with the latest data (e.g. social media firehose)?
Re: Ironwood: The first Google TPU for the age of inference
#148Re: Ironwood: The first Google TPU for the age of inference
#149Re: Ironwood: The first Google TPU for the age of inference
#150Earlier quoted context omitted.
You will own nothing and you will be happy.
Seriously yes. I don’t want to own rapidly depreciating hardware if I don’t have to. I don’t want to own open source software and I’d prefer if culture was derived from and contributed to the public domain. I’m still fully on board with capitalism, but there are many instances where I’d prefer to replace physical acquisition with renting, or replacing corporate-made culture with public culture.