Live data from Hacker News

Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

blog.google

461–470 of 616 posts

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#461

Earlier quoted context omitted.

It’s multimodal though.

> It’s multimodal though. Sure. And how does that make your day better? I know it does not improve my work in any way shape or form. I'll take a better coding model that's not multi-modal any time. If I need an LLM to do images or sound, I'd rather use a dedicated one instead of a jack-of-all-trades-master-of-none model.

Personally, I often paste screenshots into Claude Code of the application it’s working on. And I’ve even had it work autonomously on something and regularly grab its own screenshots.

Or sometimes I will have tables, charts, or even screenshots of text that I would otherwise have to have another step to OCR or type out.

Multimodal saves me time on a regular basis. Not sure it’s a game changer, but just lets me communicate with the model in all sorts of ways that would be harder otherwise.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#462
post #450

Earlier quoted context omitted.

All the benchmarks I see put it around the capabilities of Opus 4.8 Medium or Sonnet 5 High. As far as I can tell it's slightly better than GLM 5.2.

according to AA it's not better than GLM 5.2 and that's surprising to me

I personally take AA with a giant handful of salt.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#464

I wonder how big the Pro model is that Google is using behind the scenes to train these smaller ones. Going on baseless speculation, the lack of accompanying pro models with these flash releases either means: 1) the model is too big to be economical, 2) google doesn't have the compute to serve the big model, 3) their big model has too many alignment issues to serve to the public. edit: looks like benchmarks are up on…

2.5 flash was absurdly capable on a cost basis

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#465
post #107

Earlier quoted context omitted.

It's also very possible that they know their big model underperforms chatgpt 5.6 and fable by too much, so they are focusing on what they can get wins in like speed instead.

That and/or the business case isn’t as clear when serving enormous models? You’re constantly stuck in a red queen’s race where your profitability window is increasingly measured in weeks because the Chinese are right behind you. For small models (which are probably distilled from their big ones) you can serve them economically all the time and not hemorrhage money.

> For small models (which are probably distilled from their big ones) you can serve them economically all the time and not hemorrhage money.

For smaller models, you're competing with DeepSeek V4 Flash. (Which I think is a 284B A13B?) Subjectively, this feels about as smart as Sonnet 4.5, give or take. And it costs $0.09/$0.18 on Open Router, compared to $1/$5 for the latest Claude Haiku. See https://openrouter.ai/deepseek/deepseek-v4-flash#providers The developer antirez of Redis fame uses this as a local coding model.

DeepSeek did some extremely clever research on hybrid attention to get the prices that low, reducing per-user context cache sizes dramatically.

So, no, when it comes to low-price models, the US models probably can't sustain their current margins there, either.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#467
post #419

I wonder how big the Pro model is that Google is using behind the scenes to train these smaller ones. Going on baseless speculation, the lack of accompanying pro models with these flash releases either means: 1) the model is too big to be economical, 2) google doesn't have the compute to serve the big model, 3) their big model has too many alignment issues to serve to the public. edit: looks like benchmarks are up on…

I think it's 2. I frequently get told there's no capacity for Pro and the query is answered by Flash with extended thinking. And tbh it's hard to tell the difference between the two, especially if you're not coding with it.

It's hard to tell the difference because they nerfed Pro to oblivion, it used to be much, much better model (even for non-coding/chat)

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#468
post #334

Earlier quoted context omitted.

Mostly because the person I was replying to has commented about using it to write code. If you're using it for other purposes, then I give you permission to ignore my comment; there's no reason to descend into name calling.

the person you were replying to says absolutely nothing about using it to write code.

https://news.ycombinator.com/item?id=48766580

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#469
post #13

It's a bit disheartening to see no comparison to other models here - and I'm not sure this pushes the curve anywhere. 3.6 flash is more expensive than GLM 5.2 - but seemingly worse, although this post is really light (lite?) on details. It seemed for a time that Google had finally gotten the ball rolling, but I'm doubting that more and more as time passes. We'll see what happens with 3.5 pro I suppose.

My usage of GLM 5.2 has been defined by slow throughput and flaky providers, in many (definitely not all!) applications a dumber, faster, and more consistent model makes more sense to me.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#470
post #358

Earlier quoted context omitted.

In terms of open models, Gemma 4 beats the pants off everything else to the point that paying for APIs becomes hard to justify. Qwen has the meme-share for coding, but it feels much less well rounded. I have no doubt that Google have both the infrastructure and the expertise to curb stomp everyone else, should they resolve in earnest to do so. Lest we forget, "Attention is All You Need" came from Google.

Are you suggesting Gemma beats GLM 5.2?

At 20x the parameter count I should hope GLM beats Gemma! But is it 20x better? Expertise is demonstrated, not by making big models, but by making small ones. Bigger isn't better if you can't run it at all.
Post reply on HN