Earlier quoted context omitted.
The problem with publicly disclosing these is that if lots of people adopt them they will become targeted to be in the model and will no longer be a good benchmark.
This thought process is pretty baffling to me, and this is at least the second time I've encountered it on HN. What's the value of a secret benchmark to anyone but the secret holder? Does your niche benchmark even influence which model you use for unrelated queries? If LLM authors care enough about your niche (they don't) and fake the response somehow, you will learn on the very next query that something is amiss. No…
Gemini 3 Flash: Frontier intelligence built for speed
521–530 of 609 posts
Re: Gemini 3 Flash: Frontier intelligence built for speed
#522Don’t let the “flash” name fool you, this is an amazing model. I have been playing with it for the past few weeks, it’s genuinely my new favorite; it’s so fast and it has such a vast world knowledge that it’s more performant than Claude Opus 4.5 or GPT 5.2 extra high, for a fraction (basically order of magnitude less!!) of the inference time and price
[1]: https://entropicthoughts.com/haiku-4-5-playing-text-adventur...
Re: Gemini 3 Flash: Frontier intelligence built for speed
#523Earlier quoted context omitted.
Curious to learn what a “product benchmark” looks like. Is it evals you use to test prompts/models? A third party tool? Examples from the wild are a great learning tool, anything you’re able to share is appreciated.
Everyone should have their own "pelican riding a bicycle" benchmark they test new models on. And it shouldn't be shared publicly so that the models won't learn about it accidentally :)
Re: Gemini 3 Flash: Frontier intelligence built for speed
#524Earlier quoted context omitted.
Hm, quite some. Like I said, it depends what you count as AI. Just googling means you use AI nowdays.
Whether Googling something counts as AI has more to do with the shifting definition of AI over time, then with Googling itself. Remember, really back in the day the A* search algorithm was part of AI. If you had asked anyone in the 1970s whether a box that given a query pinpoints the right document that answers that question (aka Google search in the early 2000s), they'd definitely would have called it AI.
Re: Gemini 3 Flash: Frontier intelligence built for speed
#525Earlier quoted context omitted.
Not OP, but I feel the same way. Cost is just one of the factor. I'm used to Claude Code UX, my CLAUDE.md works well with my workflow too. Unless there's any significant improvement, changing to new models every few months is going to hurt me more.
just switch to Opencode and stop locking yourself into a particular providers way of doing things. There's a plugin for everything that mimics anything the others are doing
I see all of these tools as IDEs. Whether someone locks into VS Code, JetBrains, Neovim, or Sublime Text comes down to personal preference. Everyone works differently, and that is completely fine.
Re: Gemini 3 Flash: Frontier intelligence built for speed
#526Earlier quoted context omitted.
How good is it for coding, relative to recent frontier models like GPT 5.x, Sonnet 4.x, etc?
My experience so far- much less reliable. Though it’s been in chat not opencode or antigravity etc. you give it a program and say change it in this way, and it just throws stuff away, changes unrelated stuff etc. completely different quality than pro (or sonnet 4.5 / GPT-5.2)
BTW: I have the same impression, Claude was working better for me for coding tasks.
Re: Gemini 3 Flash: Frontier intelligence built for speed
#527After Gemini 3.0 the OpenAI damage control crews all drowned.
Not only is it vastly better, it's also free.
I find this particular benchmark to be in agreement with my experiences: https://simple-bench.com
Re: Gemini 3 Flash: Frontier intelligence built for speed
#528Earlier quoted context omitted.
I will have to try that. Cursor bill got pretty high with Opus 4.5. Never considered opus before the 4.5 price drop but now it's hard to change... :)
$100 Claude max is the best subscription I’ve ever had. Well worth every penny now
Re: Gemini 3 Flash: Frontier intelligence built for speed
#529Don’t let the “flash” name fool you, this is an amazing model. I have been playing with it for the past few weeks, it’s genuinely my new favorite; it’s so fast and it has such a vast world knowledge that it’s more performant than Claude Opus 4.5 or GPT 5.2 extra high, for a fraction (basically order of magnitude less!!) of the inference time and price
I wonder at what point will everyone who over-invested in OpenAI will regret their decision (expect maybe Nvidia?). Maybe Microsoft doesn't need to care, they get to sell their models via Azure.
They always had the best talent, but with Brin at the helm, they also have someone with the organizational heft to drive them towards a single goal
Re: Gemini 3 Flash: Frontier intelligence built for speed
#530Earlier quoted context omitted.
Yeah the only thing standing in Google's way is Google. And it's the easy stuff, like sensible billing models, easy to use docs and consoles that make sense and don't require 20 hours to learn/navigate, and then just the slew of bugs in Gemini CLI that are basic usability and model API interaction things. The only differentiator that OpenAI still has is polish. Edit: And just to add an example: openAI's Codex CLI bil…
I'd be curious how many people use openrouter byok just to avoid figuring out the cloud consoles for gcp/azure.