Live data from Hacker News

Please don't discontinue Gemini 2.5 Flash

discuss.ai.google.dev

31–40 of 93 posts

Re: Please don't discontinue Gemini 2.5 Flash

#31
post #18

Why not a "stop killing AI" movement? If a company deploys a paid AI model and makes people depend on it, they need to dump the weights at EOL.

Where do people get ideas like this? In what world does this make sense?

You have several choices:

1. Work with a supplier and sign a contract guaranteeing support for whatever period of time you want at a mutually agreeable price

2. Host your own stack to depend on and support it for however long you want

3. Accept that you're paying for a service and that it can go away at any time.

Companies aren't obligated to support things forever and they aren't obligated to open them up when they no longer feel it's worth supporting them. Claiming they should is absurd.

Re: Please don't discontinue Gemini 2.5 Flash

#32
Interestingly, I found the original nano banana also has the best latency/quality trade-off that new versions can't beat. This might be domain/prompt specific though. I wonder if there is some truth in the saying that something is either new or improved by never "new and improved".

Re: Please don't discontinue Gemini 2.5 Flash

#34
post #4

Can't run Qwen 3.6 35B A3B? Even Qwen 3.5 9B is comparable.

Is it really? A 9B model is equivalent? Honest question, as I haven't spent that much time with the 9B variant or Flash 2.5. But that seems like a pretty bold claim for such a small model. I assumed Flash 2.5 was considerably larger, but maybe I'm wrong?

Pretty much. It even beats it in a few benchmarks: https://artificialanalysis.ai/models/comparisons/qwen3-5-9b-...

Qwen 3.6 and Gemma 4 small models are in a league of their own.

Re: Please don't discontinue Gemini 2.5 Flash

#35
post #9

I love how there is a "Please do not discontinue gemini-2.0-flash[-lite], 2.5 is NOT an equivalent" from Feb 20th. Getting too attached to models is a smell.

the 1.5 and 2.0 flash models were absolute beasts. They were very cheap, and _very_ fast. We contemplated moving some of our fine tuned workloads to them because we would have gotten very substantial total latency reductions for our workloads.

However, they are aggressively deprecating them (OpenAI is as well), and replacing with newer models. These newer models are all reasoning models, and importantly, only bear the flash name. They are not fast. And they are very expensive!

Re: Please don't discontinue Gemini 2.5 Flash

#36
I feel this way about gpt-5-nano (EOL December 2026). It seems like the open weight models have progressed a long way since these old models were released though. Deepseek V4 Flash is even cheaper than gpt-5-nano. I'm still going to pay a cloud provider to run it for me, I'm not local inference pilled yet, but I _can_ run it myself in the future if worse comes to worst.

Objectively testable evals are one thing, but how does one judge whether a new model is adequately reproducing the subjective "writing style" of an old model that you've gotten accustomed to the feel of?

Re: Please don't discontinue Gemini 2.5 Flash

#37

I am more concerned about the cost step up from Gemini 2.5 Flash to 3.5 Flash, with the latter being roughly 3x more expensive. I thought the intention of the Flash models was to be relatively low-latency and more affordable compared to Pro, but the newer Flash models aren’t being priced as such. Then again, the era of cheap and plentiful AI might be coming to an end…

Yeah IIRC the latest Pro is $12 and Flash is $9 which is not the usual 2X-3X multiplier we see separate model grades. It also puts Flash now about 2X GLM 5.2, which is a highly capable open weight model.

Re: Please don't discontinue Gemini 2.5 Flash

#38
post #18

Why not a "stop killing AI" movement? If a company deploys a paid AI model and makes people depend on it, they need to dump the weights at EOL.

Where do people get ideas like this? In what world does this make sense? You have several choices: 1. Work with a supplier and sign a contract guaranteeing support for whatever period of time you want at a mutually agreeable price 2. Host your own stack to depend on and support it for however long you want 3. Accept that you're paying for a service and that it can go away at any time. Companies aren't obligated to su…

I mean, claiming the have a moral imperative to do it might be a bit of a stretch, but it sure would be nice - can’t blame people for wanting things.

Re: Please don't discontinue Gemini 2.5 Flash

#39
post #9

I love how there is a "Please do not discontinue gemini-2.0-flash[-lite], 2.5 is NOT an equivalent" from Feb 20th. Getting too attached to models is a smell.

We have benchmarks for our use cases, and every generation after Gemini 2.0 Flash has been a grim hit on price/performance. Costs have gone up, throughput has gone down, and performance has improved very slightly (and regressed on a few things).
Post reply on HN