Earlier quoted context omitted.
> is it a matter of absolute capacity still being insufficient for current model sizes This. Additionally, models aren't getting smaller, they are getting bigger and to be useful to a wider range of users, they also need more context to go off of, which is even more memory. Previously: https://news.ycombinator.com/item?id=42003823 It could be partially the DC, but look at the rack density... to get to an equal amount…
Thanks for the links — I went through all of them (took me a while). The point about rack density differences between SRAM-based systems like Cerebras or Groq and GPU clusters is now clear to me. What I’m still trying to understand is the economics. From this benchmark: https://artificialanalysis.ai/models/llama-4-scout/providers... Groq seems to offer near lowest prices per million tokens and the near fastest end to…
Ironwood: The first Google TPU for the age of inference
181–186 of 186 posts
Re: Ironwood: The first Google TPU for the age of inference
#182Earlier quoted context omitted.
Nvidia cloud instances have competition. What happens when you get vendor lock in with TPU and have no exit plan? Arguably it's competition that drives value creation, not capitalism.
Just don’t let yourself get stuck behind vendor lock-in, or if you do, never let yourself feel trapped, because you never are. Whenever assessing the work involved in building an integration, always assume you’ll be doing it twice. If that sounds like too much work then you shouldn’t have outsourced to begin with.
Re: Ironwood: The first Google TPU for the age of inference
#183The first specifically designed for inference? Wasn’t the original TPU inference only?
The phrasing is very precise here, it’s the first TPU for _the age of inference_, which is a novel marketing term they have defined to refer to CoT and Deep Research.
Re: Ironwood: The first Google TPU for the age of inference
#184Re: Ironwood: The first Google TPU for the age of inference
#185Earlier quoted context omitted.
Just don’t let yourself get stuck behind vendor lock-in, or if you do, never let yourself feel trapped, because you never are. Whenever assessing the work involved in building an integration, always assume you’ll be doing it twice. If that sounds like too much work then you shouldn’t have outsourced to begin with.
Switching costs are a thing everywhere in the economy. Measure twice cut once. Planning on doing the work twice is asinine.