Live data from Hacker News

Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

blog.google

441–450 of 616 posts

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#441

It's scary relying on Google's models. I have a very price sensitive workload that used to run on flash 2.5 lite - it's deprecated now. The replacement 3.1 flash lite is a lot more expensive, but now also has a sunset date. 3.5 flash lite is even more expensive. So the price is rising and you have no choice but to keep paying more and more.

And somehow, the most annoying is not even the price hike, but it is that is you expect to build a product on any of theirs models, they spend their time being deprecated and you have like to be on the lookup to start from scratch selecting a model and fitting it every year or so... Impossible to have any stability...

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#443
post #117
post #107

Earlier quoted context omitted.

It's also very possible that they know their big model underperforms chatgpt 5.6 and fable by too much, so they are focusing on what they can get wins in like speed instead.

This is the feeling i get too. Cant produce quality, but can produce something that is super fast...so take the wins where they are.

For a coding LLM specifically, when is fast a good tradeoff for quality?

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#444

Earlier quoted context omitted.

Logan Kilpatrick said on an interview not too long ago that flash 3 and 3.5 are the same pre-train. all gains on top of 3 flash are post-training

Maybe, but they said they have “started” the Gemini 4 pretrain. So not having done any significant pretrain in a year or so seems odd to me.

Pre-trains take a huge chunk of your compute offline, incurring both an raw expense (24/7 max power for all training clusters) and an opportunity cost (could have sold excess compute during that time). They also don't come with any great guarantees, as lots of techniques look good on small scale and crumble or plateau once scaled.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#445

Earlier quoted context omitted.

[flagged]

> Domestic China is the only very large audience for their own models I don't think so. US models are very expensive, and not available in every country. I am not willing to pay $50/1M tokens for writing my pet projects.

There are also US based companies like Fireworks serving up the best open weight models with the compliances we need in US enterprise. Depending on the company, they may offer more/different jurisdictions, EU probably needs a Fireworks like company (haven't heard about one, maybe it already exists?)

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#446
post #442

Here's the issue: GLM 5.2 is better, also cheaper, and almost as fast. So essentially, a big L for Google. Combine this with them not being able to produce a frontier model this generation... hmm implications

3.6 is roughly 50% faster, which isn't totally insignificant for being marginally more expensive.[1]

[1]artificialanalysis.ai

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#447
post #117

Earlier quoted context omitted.

This is the feeling i get too. Cant produce quality, but can produce something that is super fast...so take the wins where they are.

For a coding LLM specifically, when is fast a good tradeoff for quality?

I wouldnt say it is, but there are circumstances when speed is helpful. I wouldn't argue that coding is one of them.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#448
post #107

I wonder how big the Pro model is that Google is using behind the scenes to train these smaller ones. Going on baseless speculation, the lack of accompanying pro models with these flash releases either means: 1) the model is too big to be economical, 2) google doesn't have the compute to serve the big model, 3) their big model has too many alignment issues to serve to the public. edit: looks like benchmarks are up on…

It's also very possible that they know their big model underperforms chatgpt 5.6 and fable by too much, so they are focusing on what they can get wins in like speed instead.

That's the only explanation that makes sense. If it was frontier but cost or compute were limiting factors, they'd release it at an obscene price for the bragging rights. Google doesn't care that much about alignment, and I don't think it's likely to be significantly different than 3.5 anyway. The only reason it would need to be soft-canceled is if it's terrible, and has to end up in a ditch like Llama 4 to avoid shareholder panic.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#449

Earlier quoted context omitted.

[flagged]

You’re absolutely right and it’s heartening to see. I maintain a client with ~every provider you can think of and llama.cpp and it was really tiring the last few days to see people laundering other stuff through Kimi and Qwen. They’re not even open yet, the hype was based on their own blog posts, no one’s actually running these locally, the Qwen Max’s have never been open, Kimi’s API was 1/2 the speed the benchmarks…

"You’re absolutely right and it’s heartening to see"

Damnit, I usually don't jump to LLM speech patterns, but this opening had me thinking you were a bot. But after checking your profile, I think you pass as human. I wonder when will be the time, this does not work anymore for me. (Creation date is a strong hint, but abandoned accounts can be hijacked)

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#450
post #13

It's a bit disheartening to see no comparison to other models here - and I'm not sure this pushes the curve anywhere. 3.6 flash is more expensive than GLM 5.2 - but seemingly worse, although this post is really light (lite?) on details. It seemed for a time that Google had finally gotten the ball rolling, but I'm doubting that more and more as time passes. We'll see what happens with 3.5 pro I suppose.

All the benchmarks I see put it around the capabilities of Opus 4.8 Medium or Sonnet 5 High. As far as I can tell it's slightly better than GLM 5.2.

according to AA it's not better than GLM 5.2 and that's surprising to me
Post reply on HN