Live data from Hacker News

Our eighth generation TPUs: two chips for the agentic era

blog.google

121–130 of 240 posts

Re: Our eighth generation TPUs: two chips for the agentic era

#121
post #35

Which company is building the silicon for Google? Is it tsmc? What node size? I didn't see it with a quick search, sorry if it was in the post.

tsmc through broadcom

Apparently it's Broadcom for 8t and Mediatek for 8i.

https://wccftech.com/google-splits-tpuv8-strategy-two-chips-...

Re: Our eighth generation TPUs: two chips for the agentic era

#122
post #95

Earlier quoted context omitted.

> put out something really polished Like Apple Intelligence? Which was quite crap

Specifically what was crap about it? It seems to do what was advertised.

I think most people expect more than semi-reliably setting a verbal timer in 2026.

Re: Our eighth generation TPUs: two chips for the agentic era

#123

I already felt that gemini 3 proved what is possible if you train a model for efficiency. If I had to guess the pro and flash variants are 5x to 10x smaller than opus and gpt-5 class models. They produce drastically lower amount of tokens to solve a problem, but they haven't seem to have put enough effort into refinining their reasoning and execution as they produce broken toolcalls and generally struggle with 'agent…

> They produce drastically lower amount of tokens to solve a problem, but they haven't seem to have put enough effort into refinining their reasoning and execution as they produce broken toolcalls and generally struggle with 'agentic' tasks, but for raw problem solving without tools or search they match opus and gpt while presumably being a fraction of the size. Agreed, Gemini-cli is terrible compared to CC and even…

I wonder what I am missing, because I can use gemini-cli with English descriptions of features or entire projects and it just cranks out the code. Built a bunch of stuff with it. Can't think of anything it's currently lacking.

Re: Our eighth generation TPUs: two chips for the agentic era

#124
I'm surprised the interconnect per system is so slow? 6x 200Gb feels barely competitive. Same as last year.

Trainium3 and Maia 200 are 2.5 and 2.8Tb/s vs this 1.2Tb/s. Maia is 6 stacks of HBMe3, so ratio of mem:interconnect bandwidth is really falling behind here. Notably Maia is also, like TPU, high radix.

Re: Our eighth generation TPUs: two chips for the agentic era

#125

I already felt that gemini 3 proved what is possible if you train a model for efficiency. If I had to guess the pro and flash variants are 5x to 10x smaller than opus and gpt-5 class models. They produce drastically lower amount of tokens to solve a problem, but they haven't seem to have put enough effort into refinining their reasoning and execution as they produce broken toolcalls and generally struggle with 'agent…

Their "preview" naming is pretty arbitrary. It's just their way to avoid making any availability or persistence promises, let alone guarantees. It's also a PR tactic to mask any failures by pretending it's beta quality.

Re: Our eighth generation TPUs: two chips for the agentic era

#127
post #89
post #11

At this point, when you are doing big AI you basically have to buy it from NVidia or rent it from Google. And Google can design their chips and engine and systems in a whole-datacenter context, centralizing some aspects that are impossible for chip vendors to centralize, so I suspect that when things get really big, Google's systems will always be more cost-efficient. (disclosure: I am long GOOG, for this and a few o…

Isn't Amazon doing the same thing, making their own TPU's?

Yeah trainium and inferentia. They’re just not nearly as well supported on the software level. Google has already made sure this new generation will be supported by vllm, sglang, etc. Amazons chips barely support those and only multiple versions back. Super under invested in (at least on the open source side)

Re: Our eighth generation TPUs: two chips for the agentic era

#128
post #18
post #4

As others have been capturing news cycle eyes, seems to me Google has been going from strength to strength quietly in the background capturing consumer market share and without much (any?) infrastructure problems considering they're so vertically integrated in AI since day one? At one point they even seemed like a lost cause, but they're like a tide.. just growing all around.

Their latest open models are pretty competitive with other open models, and some innovation around the smaller sizes (2-4 GB). They're helping close to the distance to realistic quality inference on phones and other smaller devices.

> They're helping close to the distance to realistic quality inference on phones and other smaller devices.

If someone monopolized OS marketshare for mid- to low-priced devices, that does seem like it would be a useful research focus.

Whereas offering the same with compute-inefficiency cloud inference would be economically unviable at scale.

Free on-device Google premium closed-source models* = free Google Maps 2.0

* As long as you ship Google Apps and Play Services

Re: Our eighth generation TPUs: two chips for the agentic era

#129
post #123

Earlier quoted context omitted.

> They produce drastically lower amount of tokens to solve a problem, but they haven't seem to have put enough effort into refinining their reasoning and execution as they produce broken toolcalls and generally struggle with 'agentic' tasks, but for raw problem solving without tools or search they match opus and gpt while presumably being a fraction of the size. Agreed, Gemini-cli is terrible compared to CC and even…

I wonder what I am missing, because I can use gemini-cli with English descriptions of features or entire projects and it just cranks out the code. Built a bunch of stuff with it. Can't think of anything it's currently lacking.

>> Can't think of anything it's currently lacking.

Speed? The pro models are slow for me

The model 3.1 pro model is good and i don't recognise the GP's complaint of broken tool calls but i'm only using via gemini cli harness, sounds like they might be hosting their own agentic loop?

Re: Our eighth generation TPUs: two chips for the agentic era

#130

I'm surprised the interconnect per system is so slow? 6x 200Gb feels barely competitive. Same as last year. Trainium3 and Maia 200 are 2.5 and 2.8Tb/s vs this 1.2Tb/s. Maia is 6 stacks of HBMe3, so ratio of mem:interconnect bandwidth is really falling behind here. Notably Maia is also, like TPU, high radix.

Isn't it 6x 200Gb octals? An octal being 8x 200Gb lanes. So 9.6Tbps?
Post reply on HN