Live data from Hacker News

Improved Gemini 2.5 Flash and Flash-Lite

developers.googleblog.com

141–150 of 285 posts

Re: Improved Gemini 2.5 Flash and Flash-Lite

#141

This really captures something I've been experiencing with Gemini lately. The models are genuinely capable when they work properly, but there's this persistent truncation issue that makes them unreliable in practice. I've been running into it consistently, responses that just stop mid-sentence, not because of token limits or content filters, but what appears to be a bug in how the model signals completion. It's been…

Small things like this or the fact that AI studio still has issues with simple scrolling confuse me. How does such a brilliant tool still lack such basic things?

Re: Improved Gemini 2.5 Flash and Flash-Lite

#142

Non-AI Summary: Both models have improved intelligence on Artificial Analysis index with lower end-to-end response time. Also 24% to 50% improved output token efficiency (resulting in lower cost). Gemini 2.5 Flash-Lite improvements include better instruction following, reduced verbosity, stronger multimodal & translation capabilities. Gemini 2.5 Flash improvements include better agentic tool use and more token-effici…

I think “Non-AI summary” is going to become a thing. I already enjoyed reading it more because I knew someone had thought about the content.

As soon as it becomes a thing LLMs will start putting "Non-AI summary" at the top of their responses.

Re: Improved Gemini 2.5 Flash and Flash-Lite

#143

The switch by Artificial Analysis from per-token-cost to per-benchmark-cost shows some effect! Its nice that labs are now trying to optimize what I actually have to pay to get an answer - It always annoys me to have to pay for all the senseless rambling of the less-capable reasoning models.

Did they? I'm looking at the Artificial Analysis leaderboard site now and I only see price as USD/1M tokens.

Re: Improved Gemini 2.5 Flash and Flash-Lite

#144
post #39

Earlier quoted context omitted.

2.5 isn't the version number, its the model generation. it would only be updated when the underlying model architecture, training, etc are updated. this release is, as the name implies, the same model but likely with hardware optimizations, system prompt, and fine-tuning tweaks applied.

Ok, so if not 2.6 then 2.5.1 :)

It's model=2.5 weights=202509

Re: Improved Gemini 2.5 Flash and Flash-Lite

#146

This really captures something I've been experiencing with Gemini lately. The models are genuinely capable when they work properly, but there's this persistent truncation issue that makes them unreliable in practice. I've been running into it consistently, responses that just stop mid-sentence, not because of token limits or content filters, but what appears to be a bug in how the model signals completion. It's been…

FWIW, I think GLM-4.5 or Kimi K2 0905 fit the bill pretty well in terms of complete and consistent.

(Disclosure: I'm the founder of Synthetic.new, a company that runs open-source LLMs for monthly subscriptions.)

Re: Improved Gemini 2.5 Flash and Flash-Lite

#147
post #141

This really captures something I've been experiencing with Gemini lately. The models are genuinely capable when they work properly, but there's this persistent truncation issue that makes them unreliable in practice. I've been running into it consistently, responses that just stop mid-sentence, not because of token limits or content filters, but what appears to be a bug in how the model signals completion. It's been…

Small things like this or the fact that AI studio still has issues with simple scrolling confuse me. How does such a brilliant tool still lack such basic things?

I see Gemini web frequently break its own syntax highlighting.

Re: Improved Gemini 2.5 Flash and Flash-Lite

#148
post #109

Earlier quoted context omitted.

Unfortunately Gemini isn't the only culprit here. I've had major problems with ChatGPT reliability myself.

I only hit that problem in voice mode, it'll just stop halfway and restart. It's a jarring reminder of its lack of "real" intelligence

I've heard a lot that voice mode uses a faster (and worse) model than regular ChatGPT. So I think this makes sense. But I haven't seen this in any official documentation.

Re: Improved Gemini 2.5 Flash and Flash-Lite

#150
post #109

This really captures something I've been experiencing with Gemini lately. The models are genuinely capable when they work properly, but there's this persistent truncation issue that makes them unreliable in practice. I've been running into it consistently, responses that just stop mid-sentence, not because of token limits or content filters, but what appears to be a bug in how the model signals completion. It's been…

Unfortunately Gemini isn't the only culprit here. I've had major problems with ChatGPT reliability myself.

I think what I am seeing from ChatGPT is highly varying performance. I think this must be something they are doing to manage limitations of compute or costs. With Gemini, I think what I see is slightly different - more like a lower “peak capability” than ChatGPT’s “peak capability”.
Post reply on HN