Live data from Hacker News

Gemini 3 Flash: Frontier intelligence built for speed

blog.google

321–330 of 609 posts

Re: Gemini 3 Flash: Frontier intelligence built for speed

#321
post #67

Earlier quoted context omitted.

Calculating price increases is made more complex by the difference in token usage. From https://blog.google/products/gemini/gemini-3-flash/ : > Gemini 3 Flash is able to modulate how much it thinks. It may think longer for more complex use cases, but it also uses 30% fewer tokens on average than 2.5 Pro.

Yes, but also most of the increase in 3 Flash is in the input context price, which isn't affected by reasoning.

It is affected if it has to round-trip, e.g. because it's making tool calls.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#322

Earlier quoted context omitted.

Hi. I am curious what was the benchmark question? Cheers!

The problem with publicly disclosing these is that if lots of people adopt them they will become targeted to be in the model and will no longer be a good benchmark.

Yeah, that's part of why I don't disclose.

Obviously, the fact that I've done Google searches and tested the models on these means that their systems may have picked up on them; I'm sure that Google uses its huge dataset of Google searches and search index as inputs to its training, so Google has an advantage here. But, well, that might be why Googles new models are so much better, they're actually taking advantage of some of this massive dataset they've had for years.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#323
>"Gemini 3 Flash demonstrates that speed and scale don’t have to come at the cost of intelligence."

I am playing with Gemini 3 and the more I do the more I find it disappointing when discussing both tech and non-tech subject comparatively to ChatGPT. When it comes to non tech it seems like it was heavily indoctrinated and when it can not "prove" the point it abruptly cuts the conversation. When asked why, it says: formatting issues. Did it attend weasel courses?

It is fast. I grant it.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#324
post #204

Earlier quoted context omitted.

I'm actually liking 5.2 in Codex. It's able to take my instructions, do a good job at planning out the implementation, and will ask me relevant questions around interactions and functionality. It also gives me more tokens than Claude for the same price. Now, I'm trying to white label something that I made in Figma so my use case is a lot different from the average person on this site, but so far it's my go to and I d…

I've noticed when it comes to evaluating AI models, most people simply don't ask difficult enough questions. So everything is good enough, and the preference comes down to speed and style. It's when it becomes difficult, like in the coding case that you mentioned, that we can see the OpenAI still has the lead. The same is true for the image model, prompt adherence is significantly better than Nano Banana. Especially…

I'm currently working on a Lojban parser written in Haskell. This is a fairly complex task that requires a lot of reasoning. And I tried out all the SOTA agents extensively to see which one works the best. And Opus 4.5 is running circles around GPT-5.2 for this. So no, I don't think it's true that OpenAI "still has the lead" in general. Just in some specific tasks.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#325
post #4

Don’t let the “flash” name fool you, this is an amazing model. I have been playing with it for the past few weeks, it’s genuinely my new favorite; it’s so fast and it has such a vast world knowledge that it’s more performant than Claude Opus 4.5 or GPT 5.2 extra high, for a fraction (basically order of magnitude less!!) of the inference time and price

Alright so we have more benchmarks including hallucinations and flash doesn't do well with that, though generally it beats gemini 3 pro and GPT 5.1 thinking and gpt 5.2 thinking xhigh (but then, sonnet, grok, opus, gemini and 5.1 beat 5.2 xhigh) - everything. Crazy. https://artificialanalysis.ai/evaluations/omniscience

On your Omniscience-Index vs. Cost graph, I think your Gemini 3 pro & flash models might be swapped.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#326

Earlier quoted context omitted.

Coding is basically an edge case for LLMs too. Pretty much every person in the first (and second) world is using AI now, and only small fraction of those people are writing software. This is also reflected in OAI's report from a few months ago that found programming to only be 4% of tokens.

> Pretty much every person in the first (and second) world is using AI now This sounds like you live in a huge echo chamber. :-(

Depends what you count as AI (just googling makes you use the LLM summary), but also my mother who is really not tech affine loved what google lense can do, after I showed her.

Apart from my very old grandmothers, I don't know anyone not using AI.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#327

Earlier quoted context omitted.

Is there a "good enough" endgame for LLMs and AI where benchmarks stop mattering because end users don't notice or care? In such a scenario brand would matter more than the best tech, and OpenAI is way out in front in brand recognition.

Google biggest advantage over time will be costs. They have their own hardware which they can and will optimise for their LLMS. And Google has experience of getting market share over time by giving better results, performance or space. ie gmail vs hotmail/yahoo. Chrome vs IE/Firefox. So don't discount them if the quality is better they will get ahead over time.

It already is costs. Their Pro plan has much more generous limits compared to both OpenAI and especially Anthropic. You get 20 Deep Research queries with Pro per day, for example.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#328
post #267

Earlier quoted context omitted.

I'm a significant genAI skeptic. I periodically ask them questions about topics that are subtle or tricky, and somewhat niche, that I know a lot about, and find that they frequently provide extremely bad answers. There have been improvements on some topics, but there's one benchmark question that I have that just about every model I've tried has completely gotten wrong. Tried it on LMArena recently, got a comparison…

Hi. I am curious what was the benchmark question? Cheers!

If they told you, it would be picked up in a future model's training run.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#329
post #28

Earlier quoted context omitted.

Yes, but this Flash is a lot more powerful - beating Gemini 3 Pro on some benchmarks (and pretty close on others). I don't view this as a "new Flash" but as "a much cheaper Gemini 3 Pro/GPT-5.2"

I would be less salty if they gave us 3 Flash Lite at same price as 2.5 Flash or cheaper with better capability, but they still focus on the pricier models :(

We'll probably get 3 Flash Lite eventually, it just takes time to distill the models, and you want to start with the one that is likely to bring in more money.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#330

Earlier quoted context omitted.

Just to point this out: many of these frontier models cost isn't that far away from two orders of magnitude more than what DeepSeek charges. It doesn't compare the same, no, but with coaxing I find it to be a pretty capable competent coding model & capable of answering a lot of general queries pretty satisfactorily (but if it's a short session, why economize?). $0.28/m in, $0.42/m out. Opus 4.5 is $5/$25 (17x/60x). I…

I struggle to see the incentive to do this, I have similar thoughts for locally run models. It's only use case I can imagine is small jobs at scale perhaps something like auto complete integrated into your deployed application, or for extreme privacy, honouring NDA's etc. Otherwise, if it's a short prompt or answer, SOTA (state of the art) model will be cheap anyway and id it's a long prompt/answer, it's way more lik…

"or for extreme privacy"

Or for any privacy/IP protection at all? There is zero privacy, when using cloud based LLM models.

Post reply on HN