Live data from Hacker News

Improved Gemini 2.5 Flash and Flash-Lite

developers.googleblog.com

231–240 of 285 posts

Re: Improved Gemini 2.5 Flash and Flash-Lite

#231
post #221

Earlier quoted context omitted.

grok-4-fast is a phenomenal agentic model, and gemini flash is great for deep research leaf nodes since it's so cheap, you can segment your context a lot more than you would for pro to ensure it surfaces anything that might be valuable.

why use grok? It seems like it's constantly being throttled in order to appear more right-wing

It’s actually not. Most of the time if you ask it about a contentious political issue it will either give you a balanced view or a left-leaning one. Try it and see for yourself.

Re: Improved Gemini 2.5 Flash and Flash-Lite

#232
post #5

Gemini 2.5 Flash is an impressive model for its price. However, I don't understand why Gemini 2.0 Flash is still popular. From OpenRouter last week: * xAI: Grok Code Fast 1: 1.15T * Anthropic: Claude Sonnet 4: 586B * Google: Gemini 2.5 Flash: 325B * Sonoma Sky Alpha: 227B * Google: Gemini 2.0 Flash: 187B * DeepSeek: DeepSeek V3.1 (free): 180B * xAI: Grok 4 Fast (free): 158B * OpenAI: GPT-4.1 Mini: 157B * DeepSeek: De…

Why is Grok so popular

it was free

Re: Improved Gemini 2.5 Flash and Flash-Lite

#233

Google seems to be the main foundation model provider that's really focusing on the latency/TPS/cost dimensions. Anthropic/OpenAI are really making strides in model intelligence, but underneath some critical threshold of performance, the really long thinking times make workflows feel a lot worse in collaboration-style tools, vs a much snappier but slightly less intelligent model. It's a delicate balance, because thes…

The other day I heard gpt-5 was really an efficiency update

It was both efficiency and knowledge/reasoning update. GPT-5 excels at coding, it solves tasks the previous versions just could not do.

Re: Improved Gemini 2.5 Flash and Flash-Lite

#234

Google seems to be the main foundation model provider that's really focusing on the latency/TPS/cost dimensions. Anthropic/OpenAI are really making strides in model intelligence, but underneath some critical threshold of performance, the really long thinking times make workflows feel a lot worse in collaboration-style tools, vs a much snappier but slightly less intelligent model. It's a delicate balance, because thes…

I would be surprised if this dichotomy you're painting holds up to scrutiny. My understanding is Gemini is not far behind on "intelligence", certainly not in a way that leaves obvious doubt over where they will be over the next iteration/model cycles, where I would expect them to at least continue closing the gap. I'd be curious if you have some benchmarks to share that suggest otherwise. Meanwhile, afaik something G…

And yet my smart speakers with the Google assistant still default to a dumb model from the pre-LLM era (although my phone's version of the assistant does call Gemini). I wonder why that is, as it would be an obvious place to integrate Gemini. The bar is very very low as anything outside the standard setting alarms, checking the weather, etc. it gets wrong most of the time.

Re: Improved Gemini 2.5 Flash and Flash-Lite

#235

Serious question: If it's an improved 2.5 model, why don't they call it version 2.6? Seems annoying to have to remember if you're using the old 2.5 or the new 2.5. Kind of like when Apple released the third-gen iPad many years ago and simply called it the "new iPad" without a number.

That would be even more confusing because then it is unclear whether 2.6 Flash is better than 2.5 Pro.

Re: Improved Gemini 2.5 Flash and Flash-Lite

#236

Am I using a different Gemini from everyone else? We have Google Workspace at my job, so Gemini is baked in. It is HORRENDOUS when compared to other models. I hear a bunch of other people talking about how great Gemini is, but I've never seen it. The responses are usually either incorrect, way too long, (essays when I wanted summaries) or just...not...good. I will ask the exact same question to both Gemini and ChatGP…

It depends on what you use it for. For answering questions I tend to prefer GPT-5, but for writing (e.g. turn these informally written ideas/bullet points into a report/proposal/etc., now shorten it a bit, emphasize this idea more, etc.) it's the best by far IMHO.

Re: Improved Gemini 2.5 Flash and Flash-Lite

#237
post #226

Earlier quoted context omitted.

I would be surprised if this dichotomy you're painting holds up to scrutiny. My understanding is Gemini is not far behind on "intelligence", certainly not in a way that leaves obvious doubt over where they will be over the next iteration/model cycles, where I would expect them to at least continue closing the gap. I'd be curious if you have some benchmarks to share that suggest otherwise. Meanwhile, afaik something G…

Yesterday I asked Gemini to recalculate the timestamps of tasks in a sequence of tasks, given it's duration and the previous timestamp. It proceeded to write code which gave results like this 2025-09-26T14:32:10Z 2025-09-26T14:32:10Z200s 2025-09-26T14:32:10Z200s600s 2025-09-26T14:32:10Z200s600s300s It then proceeded to talk about how efficient this approach was for thousands of numbers. Gemini is by far the dumbest L…

They're all a little dumb. I asked claude for a python function or functions that will take in markdown in a string and return a string with ansi codes for bold, italics and underline.

It gave me a 160 line parse function.

After gaping for a short while, I implemented it in a 5 line function and a lookup table.

These vibe codes who are proud that they generated thousands of lines of code makes me wonder if they are ever reading what they generate with a critical eye.

Re: Improved Gemini 2.5 Flash and Flash-Lite

#238
post #42

Earlier quoted context omitted.

Can't agree with that. Gemini doesn't lead just on price/performance - ironically it's the best "normie" model most of the time, despite it's lack of popularity with them until very recent. It's bad at agentic stuff, especially coding. Incomparably so compared to Claude and now GPT-5. But if it's just about asking it random stuff, and especially going on for very long in the same conversation - which non-tech users h…

My pet theory without any strong foundation is because OpenAI and Anthropic have trained their models really hard to fit the sycophantic mold of: =============================== Got it — *compliment on the info you've shared*, *informal summary of task*. *Another compliment*, but *downside of question*. ---------- (relevant emoji) Bla bla bla 1. Aspect 1 2. Aspect 2 ---------- *Actual answer* ----------- (checkmark e…

Gemini does this too, but also adds a youtube link to every answer.

Just on the video link alone Gemini is making money on the free tier by pointing the hapless user at an ad while the other LLMs make zilch off the free tier.

Re: Improved Gemini 2.5 Flash and Flash-Lite

#239

Serious question: If it's an improved 2.5 model, why don't they call it version 2.6? Seems annoying to have to remember if you're using the old 2.5 or the new 2.5. Kind of like when Apple released the third-gen iPad many years ago and simply called it the "new iPad" without a number.

That would be even more confusing because then it is unclear whether 2.6 Flash is better than 2.5 Pro.

Is a 2024 Mac boo pro better than a 2025 Mac book?

Re: Improved Gemini 2.5 Flash and Flash-Lite

#240
post #145

The most annoying thing about Gemini is that it can't stop suggesting youtube videos. Even when you ask it to stop doing that, multiple times in the same conversation, it will just keep doing it.

Might be be builtin to the model because it is impossible to remove completely...

And I say this because, I added about 50 prompts in the settings to prevent video recommendations and to remove any links to videos. but I still get text saying "the linked video explains this more" even though there is no linked video.

This is not a bad way to monetise the free tier. Non of the other token providers found any way to monetise the free tier but Gemini is doing it on almost every prompt.

Post reply on HN