Live data from Hacker News

Gemini 3 Flash: Frontier intelligence built for speed

blog.google

381–390 of 609 posts

Re: Gemini 3 Flash: Frontier intelligence built for speed

#381

Earlier quoted context omitted.

The problem with publicly disclosing these is that if lots of people adopt them they will become targeted to be in the model and will no longer be a good benchmark.

This thought process is pretty baffling to me, and this is at least the second time I've encountered it on HN. What's the value of a secret benchmark to anyone but the secret holder? Does your niche benchmark even influence which model you use for unrelated queries? If LLM authors care enough about your niche (they don't) and fake the response somehow, you will learn on the very next query that something is amiss. No…

The point is that it's a litmus test for how well the models do with niche knowledge _in general_. The point isn't really to know how well the model works for that specific niche. Ideally of course you would use a few of them and aggregate the results.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#382

Earlier quoted context omitted.

Also https://artificialanalysis.ai/evaluations/omniscience Prepare to be amazed

I’m amazed by how much Gemini 3 flash hallucinates; it performs poorly in that metric (along with lots of other models). In the Hallucination Rate vs. AA-Omniscience Index chart, it’s not in the most desirable quadrant; GPT-5.1 (high), opus 4.5 and 4.5 haiku are. Can someone explain how Gemini 3 pro/flash then do so well then in the overall Omniscience: Knowledge and Hallucination Benchmark?

I'm a total noob here, but just pointing out that Omniscience Index is roughly "Accuracy - Hallucination Rate". So it simply means that their Accuracy was very high.

> In the Hallucination Rate vs. AA-Omniscience Index chart, it’s not in the most desirable quadrant

This doesn't mean much. As long as Gemini 3 has a high hallucination rate (higher than at least 50% others), it's not going to be in the most desirable quadrant by definition.

For example, let's say a model answers 99 out of 100 questions correctly. The 1 wrong answer it produces is a hallucination (i.e. confidently wrong). This amazing model would have a 100% hallucination rate as defined here, and thus not be in the most desirable quadrant. But it should still have a very high Omniscience Index.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#383
This model is breaking records on my benchmark of choice, which is 'the fraction of Hacker News comments that are positive.' Even people who avoid Google products on principle are impressed. Hardly anyone is arguing that ChatGPT is better in any respect (except brand recognition).

Re: Gemini 3 Flash: Frontier intelligence built for speed

#384
post #178

It's a cool release, but if someone on the google team reads that: flash 2.5 is awesome in terms of latency and total response time without reasoning. In quick tests this model seems to be 2x slower. So for certain use cases like quick one-token classification flash 2.5 is still the better model. Please don't stop optimizing for that!

You can still set thinking budget to 0 to completely disable reasoning, or set thinking level to minimal or low.

>You cannot disable thinking for Gemini 3 Pro. Gemini 3 Flash also does not support full thinking-off, but the minimal setting means the model likely will not think (though it still potentially can). If you don't specify a thinking level, Gemini will use the Gemini 3 models' default dynamic thinking level, "high".

https://ai.google.dev/gemini-api/docs/thinking#levels

Re: Gemini 3 Flash: Frontier intelligence built for speed

#385
post #4

Don’t let the “flash” name fool you, this is an amazing model. I have been playing with it for the past few weeks, it’s genuinely my new favorite; it’s so fast and it has such a vast world knowledge that it’s more performant than Claude Opus 4.5 or GPT 5.2 extra high, for a fraction (basically order of magnitude less!!) of the inference time and price

OpenAI made a huge mistake neglecting fast inferencing models. Their strategy was gpt 5 for everything, which hasn't worked out at all. I'm really not sure what model OpenAI wants me to use for my applications that require lower latency. If I follow their advice in their API docs about which models I should use for faster responses I get told either use GPT 5 low thinking, or replace gpt 5 with gpt 4.1, or switch to…

> OpenAI should stop trying to come up with ads and make models that are useful.

Turns out becoming a $4 trillion company first with ads (Google), then owning everybody on the AI-front could be the winning strategy.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#386
post #366
post #124

Earlier quoted context omitted.

Really stupid question: How is Gemini-like 'thinking' separate from artificial general intelligence (AGI)? When I ask Gemini 3 Flash this question, the answer is vague but agency comes up a lot. Gemini thinking is always triggered by a query. This seems like a higher-level programming issue to me. Turn it into a loop. Keep the context. Those two things make it costly for sure. But does it make it an AGI? Surely Googl…

AGI is hard but we can solve most tasks with artificial stupidity in an `until done`.

Just a matter of time and cost. Eventually...

Re: Gemini 3 Flash: Frontier intelligence built for speed

#387
post #4

Don’t let the “flash” name fool you, this is an amazing model. I have been playing with it for the past few weeks, it’s genuinely my new favorite; it’s so fast and it has such a vast world knowledge that it’s more performant than Claude Opus 4.5 or GPT 5.2 extra high, for a fraction (basically order of magnitude less!!) of the inference time and price

Oh wow - I recently tried 3 Pro preview and it was too slow for me. After reading your comment I ran my product benchmark against 2.5 flash, 2.5 pro and 3.0 flash. The results are better AND the response times have stayed the same. What an insane gain - especially considering the price compared to 2.5 Pro. I'm about to get much better results for 1/3rd of the price. Not sure what magic Google did here, but would love…

Curious to learn what a “product benchmark” looks like. Is it evals you use to test prompts/models? A third party tool?

Examples from the wild are a great learning tool, anything you’re able to share is appreciated.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#388

Earlier quoted context omitted.

I desperately want to be able to real-time dictate actions to take on my phone. Stuff like: "Open Chrome, new tab, search for xyz, scroll down, third result, copy the second paragraph, open whatsapp, hit back button, open group chat with friends, paste what we copied and send, send a follow-up laughing tears emoji, go back to chrome and close out that tab" All while being able to just quickly glance at my phone. Ther…

is that faster to say than do, or is it an accessibility or while-driving need?

[deleted]

Re: Gemini 3 Flash: Frontier intelligence built for speed

#389

Earlier quoted context omitted.

I’ve fully switched over to Gemini now. It seems significantly more useful, and is less of an automatic glaze machine that just restates your question and how smart you are for asking it.

How do I get Gemini to be more proactive in finding/double-checking itself against new world information and doing searches? For that reason I still find chatgpt way better for me, many things I ask it first goes off to do online research and has up to date information - which is surprising as you would expect Google to be way better at this. For example, was asking Gemini 3 Pro recently about how to do something wit…

Yeah any time I mention GPT-5, the other models start having panic attacks and correcting it to GPT-4. Even if it's a model name in source code!

They think GPT-5 won't be released until the distant future, but what they don't realize is we have already arrived ;)

Re: Gemini 3 Flash: Frontier intelligence built for speed

#390

This model is breaking records on my benchmark of choice, which is 'the fraction of Hacker News comments that are positive.' Even people who avoid Google products on principle are impressed. Hardly anyone is arguing that ChatGPT is better in any respect (except brand recognition).

Chatgpt 5.2 thinking is significantly better quality for most knowledge work, but it trades off in speed.
Post reply on HN