Live data from Hacker News

Gemini 3 Flash: Frontier intelligence built for speed

blog.google

561–570 of 609 posts

Re: Gemini 3 Flash: Frontier intelligence built for speed

#561
post #282
post #4

Don’t let the “flash” name fool you, this is an amazing model. I have been playing with it for the past few weeks, it’s genuinely my new favorite; it’s so fast and it has such a vast world knowledge that it’s more performant than Claude Opus 4.5 or GPT 5.2 extra high, for a fraction (basically order of magnitude less!!) of the inference time and price

I love how every single LLM model release is accompanied by pre-release insiders proclaiming how it’s the best model yet…

Thats true though.

All these announcements beat all the other models on most benchmarks and are then the best model yet. They can't see the future yet so they are not aware or care anyway that 2 weeks later someone says "hold my beer" and we get again better benchmark results from someone else.

Exhausting and exciting

Re: Gemini 3 Flash: Frontier intelligence built for speed

#563
post #544

I asked it to draft an email with a business proposal and it puts the date on letter as October 26, 2023. Then I asked it why it did so. It replies saying that the templates it was trained on might be anchored to that date. Gemini 3 Pro also puts that same date on letter. I didn't ask it why.

>ask it why

Always cracks me up asking the LLM why it said something like it really knows and won't just make up something plausible.

Scary thing is how similar we are in this regard. People confabulate and rationalize things the time, but it's especially apparent in people who engage in denial of illness (anosognosia) due to brain damage. One well documented example is stroke damaging the right hemisphere of the brain and paralyzing the left side of the body. Some will deny their paralyzed arm is paralyzed; Make up all sorts of excuses if cross examined / confronted with evidence of illness [0], or practically hallucinate their arm working, fail to notice it's not working etc. Video goes into like half a dozen experiments least. Mini spoiler: can ask someone with similar brain damage a ridiculous question "why did you just do x" (when did didn't do anything) and they'll confabulate an answer. Reminds me of split brain patients videos rationalizing why they did something (speaking left side of the brain) that was communicated visually only to the right hemisphere. [1].

Anyways, I was rewatching the anosognosia video the other day for the first time in like a decade and it really made me wonder how many evolutionary brain specializations it would take to more closely mimic human behavior in a machine.

- 0; https://www.youtube.com/watch?v=MDHJDKPeB2A - 1: https://www.youtube.com/watch?v=lfGwsAdS9Dc&t=347

Re: Gemini 3 Flash: Frontier intelligence built for speed

#564
post #523

Earlier quoted context omitted.

Any suggestions for a simple tool to set up your own local evals?

My "tool" is just prompts saved in a text file that I feed to new models by hand. I haven't built a bespoke framework on top of it. ...yet. Crap, do I need to now? =)

Yeah I’ve wondered about the same myself… My evals are also a pile of text snippets, as are some of my workflows. Thought I’d have a look to see what’s out there and found Promptfoo and Inspect AI. Haven’t tried either but will for my next round of evals

Re: Gemini 3 Flash: Frontier intelligence built for speed

#565

Earlier quoted context omitted.

How good is it for coding, relative to recent frontier models like GPT 5.x, Sonnet 4.x, etc?

My experience so far- much less reliable. Though it’s been in chat not opencode or antigravity etc. you give it a program and say change it in this way, and it just throws stuff away, changes unrelated stuff etc. completely different quality than pro (or sonnet 4.5 / GPT-5.2)

Been thinking of having Opus generate plans and then having Gemini 3 Flash execute. Might be better than using Haiku for the same.

Anyone tried something similar already?

Re: Gemini 3 Flash: Frontier intelligence built for speed

#566
post #159

Earlier quoted context omitted.

That's a quite sensationalized view. Ghibli moment was only about half a year ago. At that moment, OpenAI was so far ahead in terms of image editing. Now it's behind for a few months and "it can't be reversed"?

Check the size and budget of Google iniatives. It’s unlimited

Google basically has unlimited budget and unlimited data. If they're ahead now, which I believe they are, they'll be very very difficult to catch.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#567
post #4

Don’t let the “flash” name fool you, this is an amazing model. I have been playing with it for the past few weeks, it’s genuinely my new favorite; it’s so fast and it has such a vast world knowledge that it’s more performant than Claude Opus 4.5 or GPT 5.2 extra high, for a fraction (basically order of magnitude less!!) of the inference time and price

OpenAI made a huge mistake neglecting fast inferencing models. Their strategy was gpt 5 for everything, which hasn't worked out at all. I'm really not sure what model OpenAI wants me to use for my applications that require lower latency. If I follow their advice in their API docs about which models I should use for faster responses I get told either use GPT 5 low thinking, or replace gpt 5 with gpt 4.1, or switch to…

One can only hope OpenAI continues down the path they're on. Let them chase ads. Let them shoot themselves in the foot now. If they fail early maybe we can move beyond this ridiculous charade of generally useless models. I get it, applied in specific scenarios they have tangible use cases. But ask your non-tech caring friend or family member what frontier model was released this week and they'll not only be confused by what "frontier" means, but it's very likely they won't have any clue. Also ask them how AI is improving their lives on the daily. I'm not sure if we're at the 80% of model improvement as of yet, but given OpenAIs progress this year it seems they're at a very weak inflection point. Start serving ads so the house of cards can get a nudge.

And now with RAM, GPU and boards being a PitA to get based on supply and pricing - double middle finger to all the big tech this holiday season!

Re: Gemini 3 Flash: Frontier intelligence built for speed

#568
Curious how well it would do in Gemini CLI. Probably not that good, at least from looking at the terminal-bench-2 benchmark where it’s significantly behind Gemini-3-Pro (47.6% vs 54.2%), and I didn’t really like G3Pro in Gemini-CLI anyway. Also curious that the posted benchmark omitted comparison with Opus 4.5, which in Claude-Code is anecdotally at/near the top right now.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#569
post #384

Earlier quoted context omitted.

>You cannot disable thinking for Gemini 3 Pro. Gemini 3 Flash also does not support full thinking-off, but the minimal setting means the model likely will not think (though it still potentially can). If you don't specify a thinking level, Gemini will use the Gemini 3 models' default dynamic thinking level, "high". https://ai.google.dev/gemini-api/docs/thinking#levels

I was talking about Gemini 3 Flash, and you absolutely can disable reasoning, just try sending thinking budget: 0. It's strange that they don't want to mention this, but it works.

Gemini 3 Flash is in the second sentence.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#570
post #267
post #4

Don’t let the “flash” name fool you, this is an amazing model. I have been playing with it for the past few weeks, it’s genuinely my new favorite; it’s so fast and it has such a vast world knowledge that it’s more performant than Claude Opus 4.5 or GPT 5.2 extra high, for a fraction (basically order of magnitude less!!) of the inference time and price

I'm a significant genAI skeptic. I periodically ask them questions about topics that are subtle or tricky, and somewhat niche, that I know a lot about, and find that they frequently provide extremely bad answers. There have been improvements on some topics, but there's one benchmark question that I have that just about every model I've tried has completely gotten wrong. Tried it on LMArena recently, got a comparison…

Counter point about general knowledge that is documented/discussed in different spots on the internet.

Today I had to resolve performance problems for some sql server statement. Been doing it years, know the regular pitfalls, sometimes have to find "right" words to explain to customer why X is bad and such.

I described the issue to GPT5.2, gave the query, the execution plan and asked for help.

It was spot on, high quality responses and actionable items and explanations on why this or that is bad, how to improve it and why particularly sql may have generated such a query plan. I could instantly validate the response given my experience in the field. I even answered with some parts of chatgpt on how well it explained. However I did mention that to customer and I did tell them I approve the answer.

Asked high quality question and receive a high quality answer. And I am happy that I found out about an sql server flag where I can influence particular decision. But the suggestion was not limited to that, there were multiple points given that would help.

Post reply on HN