Earlier quoted context omitted.
Hardware is a factor here. GPUs are necessarily higher latency than TPUs for equivalent compute on equivalent data. There are lots of other factors here, but latency specifically favours TPUs. The only non-TPU fast models I'm aware of are things running on Cerebras can be much faster because of their CPUs, and Grok has a super fast mode, but they have a cheat code of ignoring guardrails and making up their own world…
Why are GPUs necessarily higher latency than TPUs? Both require roughly the same arithmetic intensity and use the same memory technology at roughly the same bandwidth.
Gemini 3 Flash: Frontier intelligence built for speed
421–430 of 609 posts
Re: Gemini 3 Flash: Frontier intelligence built for speed
#422Earlier quoted context omitted.
Hard to find info but I think the -chat versions of 5.1 and 5.2 (gpt-5.2-chat) are what you're looking for. They might just be an alias for the same model with very low reasoning though. I've seen other providers do the same thing, where they offer a reasoning and non reasoning endpoint. Seems to work well enough.
It's weird they don't document this stuff. Like understanding things like tool call latency and time to first token is extremely important in application development.
Re: Gemini 3 Flash: Frontier intelligence built for speed
#423Earlier quoted context omitted.
it's easy to comprehend actually. they're putting everything on "having the best model". It doesn't look like they're going to win, but that's still their bet/
I mean they’re trying to outdo google. So they need to do that.
Re: Gemini 3 Flash: Frontier intelligence built for speed
#424Earlier quoted context omitted.
OpenAI's doom was written when Altman (and Nadella) got greedy, threw away the nonprofit mission, and caused the exodus of talent and funding that created Anthropic. If they had stayed nonprofit the rest of the industry could have consolidated their efforts against Google's juggernaut. I don't understand how they expected to sustain the advantage against Google's infinite money machine. With Waymo Google showed that…
> I don't understand how they expected to sustain the advantage against Google's infinite money machine. I ask this question about Nazi Germany. They adopted the Blitkrieg strategy and expanded unsustainably, but it was only a matter of time until powers with infinite resources (US, USSR) put an end to it.
Not saying that the Nazi strategy was without flaws, of course. But your specific critique is a bit too blunt.
Re: Gemini 3 Flash: Frontier intelligence built for speed
#425Don’t let the “flash” name fool you, this is an amazing model. I have been playing with it for the past few weeks, it’s genuinely my new favorite; it’s so fast and it has such a vast world knowledge that it’s more performant than Claude Opus 4.5 or GPT 5.2 extra high, for a fraction (basically order of magnitude less!!) of the inference time and price
I'm a significant genAI skeptic. I periodically ask them questions about topics that are subtle or tricky, and somewhat niche, that I know a lot about, and find that they frequently provide extremely bad answers. There have been improvements on some topics, but there's one benchmark question that I have that just about every model I've tried has completely gotten wrong. Tried it on LMArena recently, got a comparison…
Re: Gemini 3 Flash: Frontier intelligence built for speed
#426Earlier quoted context omitted.
I wonder at what point will everyone who over-invested in OpenAI will regret their decision (expect maybe Nvidia?). Maybe Microsoft doesn't need to care, they get to sell their models via Azure.
Oracle's stock skyrocketed then took a nosedive. Financial experts warned that companies who bet big on OpenAI like Oracle and Coreweave to pump their stock would go down the drain, and down the drain they went (so far: -65% for Coreweave and nearly -50% of Oracle compared to their OpenAI-hype all-time highs). Markets seems to be in a: "Show me the OpenAI money" mood at the moment. And even financial commentators who…
[0] At least the guys who publish where you or me can read them.
Re: Gemini 3 Flash: Frontier intelligence built for speed
#427Earlier quoted context omitted.
I love how every single LLM model release is accompanied by pre-release insiders proclaiming how it’s the best model yet…
Make me think of how every iPhone is the best iPhone yet. Waiting for Apple to say "sorry folks, bad year for iPhone"
Re: Gemini 3 Flash: Frontier intelligence built for speed
#428This is awesome. No preview release either, which is great to production. They are pushing the prices higher with each release though: API pricing is up to $0.5/M for input and $3/M for output For comparison: Gemini 3.0 Flash: $0.50/M for input and $3.00/M for output Gemini 2.5 Flash: $0.30/M for input and $2.50/M for output Gemini 2.0 Flash: $0.15/M for input and $0.60/M for output Gemini 1.5 Flash: $0.075/M for inp…
This is a preview release.
Re: Gemini 3 Flash: Frontier intelligence built for speed
#429Earlier quoted context omitted.
"or for extreme privacy" Or for any privacy/IP protection at all? There is zero privacy, when using cloud based LLM models.
Really only if you are paranoid. It's incredibly unlikely that the labs are lying about not training on your data for the API plans that offer it. Breaking trust with outright lies would be catastrophic to any lab right now. Enterprise demands privacy, and the labs will be happy to accommodate (for the extra cost, of course).
Re: Gemini 3 Flash: Frontier intelligence built for speed
#430Earlier quoted context omitted.
There will be diminishing returns though as the future models won't be thah much better we will reach a point where the open source model will be good enough for most things. And the need for being on the latest model no longer so important. For me the bigger concern which I have mentioned on other AI related topics is that AI is eating all the production of computer hardware so we should be worrying about hardware p…
I had a similar opinion, that we were somewhere near the top of the sigmoid curve of model improvement that we could achieve in the near term. But given continued advancements, I’m less sure that prediction holds.
So I don't think we are on any sigmoid curve or so. Though if you plot the performance of the best model available at any point in time against time on the x-axis, you might see a sigmoid curve, but that's a combination of the logarithm and the amount of effort people are willing to spend on making new models.
(I'm not sure about it specifically being the logarithm. Just any curve that has rapidly diminishing marginal returns that nevertheless never go to zero, ie the curve never saturates.)