Live data from Hacker News

Our eighth generation TPUs: two chips for the agentic era

blog.google

161–170 of 240 posts

Re: Our eighth generation TPUs: two chips for the agentic era

#161
post #151

I already felt that gemini 3 proved what is possible if you train a model for efficiency. If I had to guess the pro and flash variants are 5x to 10x smaller than opus and gpt-5 class models. They produce drastically lower amount of tokens to solve a problem, but they haven't seem to have put enough effort into refinining their reasoning and execution as they produce broken toolcalls and generally struggle with 'agent…

> If I had to guess the pro and flash variants are 5x to 10x smaller than opus and gpt-5 class models. I really doubt it, especially Pro. If anything I wouldn't be surprised if their hardware lets them run bigger models more cheaply and quickly than the others. Pro is probably smaller than GPT 5.4 and Opus 4.6 (looks like 4.7 decreased in size), but 5x seems way too much. IMO Gemini 3 Pro is the most "intelligent" in…

Aistudio should be their default app

Re: Our eighth generation TPUs: two chips for the agentic era

#162
post #151

I already felt that gemini 3 proved what is possible if you train a model for efficiency. If I had to guess the pro and flash variants are 5x to 10x smaller than opus and gpt-5 class models. They produce drastically lower amount of tokens to solve a problem, but they haven't seem to have put enough effort into refinining their reasoning and execution as they produce broken toolcalls and generally struggle with 'agent…

> If I had to guess the pro and flash variants are 5x to 10x smaller than opus and gpt-5 class models. I really doubt it, especially Pro. If anything I wouldn't be surprised if their hardware lets them run bigger models more cheaply and quickly than the others. Pro is probably smaller than GPT 5.4 and Opus 4.6 (looks like 4.7 decreased in size), but 5x seems way too much. IMO Gemini 3 Pro is the most "intelligent" in…

Regarding Anthropic, they used to make best multilingual and generalist models, it's their policy thing, not a capability issue. Claude 3 was best at this, including dead and low-resource languages. Neither modern Claude nor Gemini are remotely close to what Claude 3 was capable of (e.g. zero-shot writing styles). Anthropic basically reversed their "character training" policy and started optimizing their models for code generation at the cost of everything else, starting with Sonnet 3.5. Claude 4 took a huge hit in multilingual ability

GPT, on the other hand, was always terrible at languages, except for the short-lived gpt-4.5-preview.

All modern models including Gemini have bugs in basic language coherency - random language switching, self-correction attempts resulting in hallucinations etc. I speculate it's a problem with heavy RL with rewards and policies not optimized for creative writing.

Re: Our eighth generation TPUs: two chips for the agentic era

#163
post #56

In recent discussions about Tim Apple [sic] moving on there was a discussion about whether Apple flopped on AI, which is my opinion. Of course you had the false dichotomy of doing nothing or burning money faster than the US military like OpenAI does. IMHO that happy medium is Google. Not having to pay the NVidia tax will likely be a huge competitive advantage. And nobody builds data centers as cost-effectively as Goo…

Apple has not flopped on AI as you say. They are just focused on privacy and are likely waiting for the time when local models become efficient enough to run on iPhones (which is quickly becoming a reality). Google could probably train models for orders of magnitude less money as you say, but they aren't. They are not capable of creating high quality models like OpenAI and Anthropic are. Their company is just too dis…

> hey are just focused on privacy and are likely waiting for the time when local models become efficient enough to run on iPhones (which is quickly becoming a reality).

This is such revisionist history. They were not strategicially waiting. They tried, really really hard. The entire iPhone 16 pro was built on AI. Heck, they even (re)named it as Apple Intelligence.

Remember, this is the same time when Microsoft launched Copilot (RIP), Google launched Gemini, OpenAI with ChatGPT etc.

--- They had to walk back hard because it was a flop. They might be accidentally successful because they are a company with multiple strengths, but dont think of it as they were sitting AI out.

Re: Our eighth generation TPUs: two chips for the agentic era

#164

I already felt that gemini 3 proved what is possible if you train a model for efficiency. If I had to guess the pro and flash variants are 5x to 10x smaller than opus and gpt-5 class models. They produce drastically lower amount of tokens to solve a problem, but they haven't seem to have put enough effort into refinining their reasoning and execution as they produce broken toolcalls and generally struggle with 'agent…

> a model that will be an entire generation beyond SOTA

That model would then be SOTA.

Tautologically you can't be better than SOTA

Re: Our eighth generation TPUs: two chips for the agentic era

#165

Earlier quoted context omitted.

Is the internet bigger or smaller than it was in 1998 compared to today? Demand for internet and web services is significantly higher today than in 2000 but a bubble still popped. Heck a regular old recession or depression, completely unrelated to AI could happen next year and could collapse the industry. I mean housing is more expensive than ever nearly 20 years after collapsing in the Great Recession.

The problem that I have with dotcom comparisons is that people miss what popped and what remained after that bubble. Catsdotcom and Dogsdotcom popped. But the tech remained, and now we have FAANG++. If we apply the same logic, any of oAI, xAI, Anthropic might pop, but realistically they won't, and even if they do, some other players will take their spots, and the tech will survive, and more importantly the demand wil…

I feel like this fails to address my point.

In 2008 there was a subprime mortgage crisis that caused the housing market to crash. Nearly all banks who participated in this survived. There was and still is significant demand for houses, financed through mortgages.

The bubble can burst, most if not all the big players still survive 20 years later and yet significant value and capital can still be destroyed in the process.

Same for the dot com. There was demand for the internet, it couldn’t meet the expectations of the day, and yet here we are with like 100x more internet services than before all these years later. Saying the AI bubble will pop is not a prediction that all AI companies will cease to exist immediately. Amazon lost 80% of their stock price in 2000. Is Amazon bigger or smaller than they were in 2000 today?

Re: Our eighth generation TPUs: two chips for the agentic era

#166
post #4

As others have been capturing news cycle eyes, seems to me Google has been going from strength to strength quietly in the background capturing consumer market share and without much (any?) infrastructure problems considering they're so vertically integrated in AI since day one? At one point they even seemed like a lost cause, but they're like a tide.. just growing all around.

> seems to me Google has been going from strength to strength quietly in the background capturing consumer market share and without much (any?) infrastructure problems considering they're so vertically integrated in AI since day one? The Google Antigravity subreddit is a shitshow though: https://www.reddit.com/r/GoogleAntigravityIDE/

can't let google have a success without it getting autodestructed though. classic google

Re: Our eighth generation TPUs: two chips for the agentic era

#167

Earlier quoted context omitted.

They have to have SOME competitive advantage. What reason is there to use Gemini over Claude or ChatGPT? It's not producing nearly the quality of output.

I recently did my taxes using all three models (My return is ~50 pages, much more than a standard 1040). GPT (codex) was accurate on the first run and took 12 minutes Gemini (antigravity) missed 1 value because it didn't load the full 1099 pdf (the laziness), but corrected it when prompted. However it only spent 2 minutes on the task. Claude (CC) made all manner of mistakes after waiting overnight for it to finish be…

Yep, I've found Gemini to be the best LLM at most tasks that are not coding. Sometimes Opus wins for engineering, but Gemini holds its own there as well. I also used Gemini to assist me with understanding the details of my (pre-revenue) C-Corp taxes this year. It did a pretty good job walking me through each question I had and raising concern about things I might have overlooked. I validated everything against reliable sources, of course.

Gemini missed on some nuances about the paperwork processes of Delaware. Gemini repeatedly assumed I could do something instantly via an online portal that actually required either snail-mail or the use of an intermediate who actually had API access to Delaware's systems. In the end, these processes took a couple days, and while I got things done in time, I wish I had not taken questions of process at face value, and instead wish I had kicked off the taxes at the end of February rather than week before they were due.

Re: Our eighth generation TPUs: two chips for the agentic era

#168
post #16

The pics of the cooling system is pretty good sci-fi / cyberpunk / steampunk inspo. If the whole AI bubble spectularly collapes, at least we got a lot of cool pics of custom hardware!

> If the whole AI bubble spectularly collapes Every other news for the past month has been about lacking capacity. Everyone is having scaling issues with more demand than they can cover. Anthropic has been struggling for a few months, especially visible when EU tz is still up and US east coast comes online. Everything grinds to a halt. MS has been pausing new subscriptions for gh Copilot, also because a lack of capac…

Username checks out

Re: Our eighth generation TPUs: two chips for the agentic era

#169

Earlier quoted context omitted.

> seems to me Google has been going from strength to strength quietly in the background capturing consumer market share and without much (any?) infrastructure problems considering they're so vertically integrated in AI since day one? The Google Antigravity subreddit is a shitshow though: https://www.reddit.com/r/GoogleAntigravityIDE/

Damn, you're not kidding. Might be worse than r/ClaudeAI in terms of user sentiment, and that's saying something.

I mean, reddit is just a knob sama can turn for easy astroturfing. It's almost as bad as looking for grok sentiment on X.

Re: Our eighth generation TPUs: two chips for the agentic era

#170

I already felt that gemini 3 proved what is possible if you train a model for efficiency. If I had to guess the pro and flash variants are 5x to 10x smaller than opus and gpt-5 class models. They produce drastically lower amount of tokens to solve a problem, but they haven't seem to have put enough effort into refinining their reasoning and execution as they produce broken toolcalls and generally struggle with 'agent…

I really wonder what I’m missing with Gemini. It’s a second rate model for me at best. I find it okay (not great) at collecting information and completely useless at agentic tasks. It’s like it’s always drunk. When the Claude credits expire in Antigravity, I’m done for the day.

> They produce drastically lower amount of tokens to solve a problem

I LOLed at this because I of the constant death loops that don’t even solve the problem at all.

Post reply on HN