Earlier quoted context omitted.
Training their own, closed, internal models on their own data sets? Probably a good way to squeeze out some market trading signals.
I thought these TPUs were primarily used for inference?
Our eighth generation TPUs: two chips for the agentic era
51–60 of 240 posts
Re: Our eighth generation TPUs: two chips for the agentic era
#52Earlier quoted context omitted.
I'd bet that too if their management wasn't so incredibly uninspiring. Like, Apple under Cook was also pretty mild and a huge step down from Jobs, but Google feels like it fell off a cliff. If it wasn't for OpenAI releasing ChatGPT, they might still be sitting on that tech while only testing it internally. Now it drives their entire chip R&D.
What would an inspiring leader do differently for you?
Re: Our eighth generation TPUs: two chips for the agentic era
#53Re: Our eighth generation TPUs: two chips for the agentic era
#54Earlier quoted context omitted.
That's because these mega monopolies have diverse income streams and have grown like cancers to tax every system and economy that touches the internet. Anthropic and OpenAI are having to fight like hell to secure market share. Google just gets to sit back and relax with its browser and android monopolies. Why did our regulators fall asleep at the wheel? Google owns 92% of "URL bar" surface area and turned it into a G…
Google invented the AI architecture that Anthropic and OpenAI based their entire companies on? Based off years of research at Google. Of course they should have to fight with the inventors of the technology they’re using.
Source?
Re: Our eighth generation TPUs: two chips for the agentic era
#55Re: Our eighth generation TPUs: two chips for the agentic era
#56IMHO that happy medium is Google. Not having to pay the NVidia tax will likely be a huge competitive advantage. And nobody builds data centers as cost-effectively as Google. It's kind of crazy to be talking ExaFLOPS and Tb/s here. From some quick Googling:
- The first MegaFLOPS CPU was in 1964
- A Cray supercomputer hit GigaFLOPS in 1988 with workstations hitting it in the 1990s. Consumer CPUs I think hit this around 1999 with the Pentium 3 at 1GHz+;
- It was the 2010s before we saw off-the-shelf TFLOPS;
- It was only last year where a single chip hit PetaFLOPS. I see the IBM Roadrunner hit this in 2008 but that was ~13,000 CPUs so...
Obviously this is near 10,000 TPUs to get to ~121 EFLOPS (FP4 admittedly) but that's still an astounding number. IT means each one is doing ~12 PFLOPS (FP4).
I saw a claim that Claude Mythos cost ~$10B to train. I personally believe Google can (or soon will be able to) do this for an order of magnitude less at least.
I would love to know the true cost/token of Claude, ChatGPT and Gemini. I think you'll find Google has a massive cost advantage here.
Re: Our eighth generation TPUs: two chips for the agentic era
#57> A single TPU 8t superpod now scales to 9,600 chips and two petabytes of shared high bandwidth memory, with double the interchip bandwidth of the previous generation. This architecture delivers 121 ExaFlops of compute and allows the most complex models to leverage a single, massive pool of memory. This seems impressive. I don't know much about the space, so maybe it's not actually that great, but from my POV it look…
Re: Our eighth generation TPUs: two chips for the agentic era
#58At this point, when you are doing big AI you basically have to buy it from NVidia or rent it from Google. And Google can design their chips and engine and systems in a whole-datacenter context, centralizing some aspects that are impossible for chip vendors to centralize, so I suspect that when things get really big, Google's systems will always be more cost-efficient. (disclosure: I am long GOOG, for this and a few o…
I'd go long Google too if using Gemini CLI felt anything close to the experience I get with Codex or Claude. They might have great hardware but it's worthless if their flagship coding agent gets stuck in loops trying to find the end of turn token.
Re: Our eighth generation TPUs: two chips for the agentic era
#59Whats interesting to note, as someone who uses Gemini, ChatGPT, and Claude, is that Gemini consistently uses drastically fewer tokens than the other two. It seems like gemini is where it is because it has a much smaller thinking budget. It's hard to reconcile this because Google likely has the most compute and at the lowest cost, so why aren't they gassing the hell out of inference compute like the other two? Maybe a…
I was planning on comparing them on coding but I didn't get the Gemini VSCode add-in to work so yeah, no dice.
The Android and web app is also riddled with bugs, including ones that makes you lose your chat history from the threads if you switch between them, not cool.
I'll be cancelling my Google One subscription this month.
Re: Our eighth generation TPUs: two chips for the agentic era
#60Whats interesting to note, as someone who uses Gemini, ChatGPT, and Claude, is that Gemini consistently uses drastically fewer tokens than the other two. It seems like gemini is where it is because it has a much smaller thinking budget. It's hard to reconcile this because Google likely has the most compute and at the lowest cost, so why aren't they gassing the hell out of inference compute like the other two? Maybe a…
They have to have SOME competitive advantage. What reason is there to use Gemini over Claude or ChatGPT? It's not producing nearly the quality of output.