Live data from Hacker News

Ironwood: The first Google TPU for the age of inference

blog.google

181–186 of 186 posts

Re: Ironwood: The first Google TPU for the age of inference

#181
post #180

Earlier quoted context omitted.

> is it a matter of absolute capacity still being insufficient for current model sizes This. Additionally, models aren't getting smaller, they are getting bigger and to be useful to a wider range of users, they also need more context to go off of, which is even more memory. Previously: https://news.ycombinator.com/item?id=42003823 It could be partially the DC, but look at the rack density... to get to an equal amount…

Thanks for the links — I went through all of them (took me a while). The point about rack density differences between SRAM-based systems like Cerebras or Groq and GPU clusters is now clear to me. What I’m still trying to understand is the economics. From this benchmark: https://artificialanalysis.ai/models/llama-4-scout/providers... Groq seems to offer near lowest prices per million tokens and the near fastest end to…

Loss leader. It is uber/airbnb. Book revenue, regardless of economics, and then debt finance against that. Hope one day to lock in customers, or raise prices, or sell the company.

Re: Ironwood: The first Google TPU for the age of inference

#182

Earlier quoted context omitted.

Nvidia cloud instances have competition. What happens when you get vendor lock in with TPU and have no exit plan? Arguably it's competition that drives value creation, not capitalism.

Just don’t let yourself get stuck behind vendor lock-in, or if you do, never let yourself feel trapped, because you never are. Whenever assessing the work involved in building an integration, always assume you’ll be doing it twice. If that sounds like too much work then you shouldn’t have outsourced to begin with.

Switching costs are a thing everywhere in the economy. Measure twice cut once. Planning on doing the work twice is asinine.

Re: Ironwood: The first Google TPU for the age of inference

#183
post #12

The first specifically designed for inference? Wasn’t the original TPU inference only?

The phrasing is very precise here, it’s the first TPU for _the age of inference_, which is a novel marketing term they have defined to refer to CoT and Deep Research.

As a previous boss liked to say, this car is the cheapest in its price range, the roomiest in it's size category, and the fastest in its speed group.

Re: Ironwood: The first Google TPU for the age of inference

#185

Earlier quoted context omitted.

Just don’t let yourself get stuck behind vendor lock-in, or if you do, never let yourself feel trapped, because you never are. Whenever assessing the work involved in building an integration, always assume you’ll be doing it twice. If that sounds like too much work then you shouldn’t have outsourced to begin with.

Switching costs are a thing everywhere in the economy. Measure twice cut once. Planning on doing the work twice is asinine.

Then you shouldn’t have outsourced to begin with.
Post reply on HN