These flash models keep getting more expensive with every release. Is there an OSS model that's better than 2.0 flash with similar pricing, speed and a 1m context window? Edit: this is not the typical flash model, it's actually an insane value if the benchmarks match real world usage. > Gemini 3 Flash achieves a score of 78%, outperforming not only the 2.5 series, but also Gemini 3 Pro. It strikes an ideal balance fo…
cost of e2e task resolution should be cheaper, even if single inference cost is higher, you need fewer loops to solve a problem now
Gemini 3 Flash: Frontier intelligence built for speed
21–30 of 609 posts
Re: Gemini 3 Flash: Frontier intelligence built for speed
#22Deepmind Page: https://deepmind.google/models/gemini/flash/ Developer Blog: https://blog.google/technology/developers/build-with-gemini-... Model Card [pdf]: https://deepmind.google/models/model-cards/gemini-3-flash/ Gemini 3 Flash in Search AI mode: https://blog.google/products/search/google-ai-mode-update-ge...
Re: Gemini 3 Flash: Frontier intelligence built for speed
#23Pipe dream right now, but 50 years later? Maybe
Re: Gemini 3 Flash: Frontier intelligence built for speed
#24These flash models keep getting more expensive with every release. Is there an OSS model that's better than 2.0 flash with similar pricing, speed and a 1m context window? Edit: this is not the typical flash model, it's actually an insane value if the benchmarks match real world usage. > Gemini 3 Flash achieves a score of 78%, outperforming not only the 2.5 series, but also Gemini 3 Pro. It strikes an ideal balance fo…
Re: Gemini 3 Flash: Frontier intelligence built for speed
#25Google keeps their models very "fresh" and I tend to get more correct answers when asking about Azure or O365 issues, ironically copilot will talk about now deleted or deprecated features more often.
Re: Gemini 3 Flash: Frontier intelligence built for speed
#26They went too far, now the Flash model is competing with their Pro version. Better SWE-bench, better ARC-AGI 2 than 3.0 Pro. I imagine they are going to improve 3.0 Pro before it's no more in Preview. Also I don't see it written in the blog post but Flash supports more granular settings for reasoning: minimal, low, medium, high (like openai models), while pro is only low and high.
> Matches the “no thinking” setting for most queries. The model may think very minimally for complex coding tasks. Minimizes latency for chat or high throughput applications.
I'd prefer a hard "no thinking" rule than what this is.
Re: Gemini 3 Flash: Frontier intelligence built for speed
#27Re: Gemini 3 Flash: Frontier intelligence built for speed
#28Yet again Flash receives a notable price hike: from $0.3/$2.5 for 2.5 Flash to $0.5/$3 (+66.7% input, +20% output) for 3 Flash. Also, as a reminder, 2 Flash used to be $0.1/$0.4.
I don't view this as a "new Flash" but as "a much cheaper Gemini 3 Pro/GPT-5.2"
Re: Gemini 3 Flash: Frontier intelligence built for speed
#29Don’t let the “flash” name fool you, this is an amazing model. I have been playing with it for the past few weeks, it’s genuinely my new favorite; it’s so fast and it has such a vast world knowledge that it’s more performant than Claude Opus 4.5 or GPT 5.2 extra high, for a fraction (basically order of magnitude less!!) of the inference time and price
Re: Gemini 3 Flash: Frontier intelligence built for speed
#30They are pushing the prices higher with each release though: API pricing is up to $0.5/M for input and $3/M for output
For comparison:
Gemini 3.0 Flash: $0.50/M for input and $3.00/M for output
Gemini 2.5 Flash: $0.30/M for input and $2.50/M for output
Gemini 2.0 Flash: $0.15/M for input and $0.60/M for output
Gemini 1.5 Flash: $0.075/M for input and $0.30/M for output (after price drop)
Gemini 3.0 Pro: $2.00/M for input and $12/M for output
Gemini 2.5 Pro: $1.25/M for input and $10/M for output
Gemini 1.5 Pro: $1.25/M for input and $5/M for output
I think image input pricing went up even more.
Correction: It is a preview model...