Live data from Hacker News

Gemini 3 Flash: Frontier intelligence built for speed

blog.google

521–530 of 609 posts

Re: Gemini 3 Flash: Frontier intelligence built for speed

#521

Earlier quoted context omitted.

The problem with publicly disclosing these is that if lots of people adopt them they will become targeted to be in the model and will no longer be a good benchmark.

This thought process is pretty baffling to me, and this is at least the second time I've encountered it on HN. What's the value of a secret benchmark to anyone but the secret holder? Does your niche benchmark even influence which model you use for unrelated queries? If LLM authors care enough about your niche (they don't) and fake the response somehow, you will learn on the very next query that something is amiss. No…

Because it encompasses the very specific way I like to do things. It's not of use to the general public.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#522
post #4

Don’t let the “flash” name fool you, this is an amazing model. I have been playing with it for the past few weeks, it’s genuinely my new favorite; it’s so fast and it has such a vast world knowledge that it’s more performant than Claude Opus 4.5 or GPT 5.2 extra high, for a fraction (basically order of magnitude less!!) of the inference time and price

Yes, 2.5 Flash is extremely cost efficient in my favourite private benchmark: playing text adventures[1]. I'm looking forward to testing 3.0 Flash later today.

[1]: https://entropicthoughts.com/haiku-4-5-playing-text-adventur...

Re: Gemini 3 Flash: Frontier intelligence built for speed

#523

Earlier quoted context omitted.

Curious to learn what a “product benchmark” looks like. Is it evals you use to test prompts/models? A third party tool? Examples from the wild are a great learning tool, anything you’re able to share is appreciated.

Everyone should have their own "pelican riding a bicycle" benchmark they test new models on. And it shouldn't be shared publicly so that the models won't learn about it accidentally :)

Any suggestions for a simple tool to set up your own local evals?

Re: Gemini 3 Flash: Frontier intelligence built for speed

#524
post #436
post #345

Earlier quoted context omitted.

Hm, quite some. Like I said, it depends what you count as AI. Just googling means you use AI nowdays.

Whether Googling something counts as AI has more to do with the shifting definition of AI over time, then with Googling itself. Remember, really back in the day the A* search algorithm was part of AI. If you had asked anyone in the 1970s whether a box that given a query pinpoints the right document that answers that question (aka Google search in the early 2000s), they'd definitely would have called it AI.

Google gives you an AI summary, reading that means interacting with LLMs.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#525
post #91

Earlier quoted context omitted.

Not OP, but I feel the same way. Cost is just one of the factor. I'm used to Claude Code UX, my CLAUDE.md works well with my workflow too. Unless there's any significant improvement, changing to new models every few months is going to hurt me more.

just switch to Opencode and stop locking yourself into a particular providers way of doing things. There's a plugin for everything that mimics anything the others are doing

Being open does not magically make everything better. People are willing to pay for Claude Code for many valid reasons. You are also assuming I have never used OpenCode, which is incorrect. Claude is simply my preference.

I see all of these tools as IDEs. Whether someone locks into VS Code, JetBrains, Neovim, or Sublime Text comes down to personal preference. Everyone works differently, and that is completely fine.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#526

Earlier quoted context omitted.

How good is it for coding, relative to recent frontier models like GPT 5.x, Sonnet 4.x, etc?

My experience so far- much less reliable. Though it’s been in chat not opencode or antigravity etc. you give it a program and say change it in this way, and it just throws stuff away, changes unrelated stuff etc. completely different quality than pro (or sonnet 4.5 / GPT-5.2)

So why Flash is so high in LiveCodeBench Pro?

BTW: I have the same impression, Claude was working better for me for coding tasks.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#528
post #106

Earlier quoted context omitted.

I will have to try that. Cursor bill got pretty high with Opus 4.5. Never considered opus before the 4.5 price drop but now it's hard to change... :)

$100 Claude max is the best subscription I’ve ever had. Well worth every penny now

Or a $40 GitHub copilot plan also gets you a lot of Opus usage.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#529
post #4

Don’t let the “flash” name fool you, this is an amazing model. I have been playing with it for the past few weeks, it’s genuinely my new favorite; it’s so fast and it has such a vast world knowledge that it’s more performant than Claude Opus 4.5 or GPT 5.2 extra high, for a fraction (basically order of magnitude less!!) of the inference time and price

I wonder at what point will everyone who over-invested in OpenAI will regret their decision (expect maybe Nvidia?). Maybe Microsoft doesn't need to care, they get to sell their models via Azure.

Seeing Sergey Brin back in the trenches makes me think Google is really going to win this

They always had the best talent, but with Brin at the helm, they also have someone with the organizational heft to drive them towards a single goal

Re: Gemini 3 Flash: Frontier intelligence built for speed

#530

Earlier quoted context omitted.

Yeah the only thing standing in Google's way is Google. And it's the easy stuff, like sensible billing models, easy to use docs and consoles that make sense and don't require 20 hours to learn/navigate, and then just the slew of bugs in Gemini CLI that are basic usability and model API interaction things. The only differentiator that OpenAI still has is polish. Edit: And just to add an example: openAI's Codex CLI bil…

I'd be curious how many people use openrouter byok just to avoid figuring out the cloud consoles for gcp/azure.

Openrouter is great! Prepaid, no surprise bills. Easily switch between any models you desire. Dead simple interface. Reliable. What's not to like?
Post reply on HN