Live data from Hacker News

Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

blog.google

471–480 of 616 posts

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#471
post #107

Earlier quoted context omitted.

It's also very possible that they know their big model underperforms chatgpt 5.6 and fable by too much, so they are focusing on what they can get wins in like speed instead.

That's the only explanation that makes sense. If it was frontier but cost or compute were limiting factors, they'd release it at an obscene price for the bragging rights. Google doesn't care that much about alignment, and I don't think it's likely to be significantly different than 3.5 anyway. The only reason it would need to be soft-canceled is if it's terrible, and has to end up in a ditch like Llama 4 to avoid sha…

3.5 pro was clearly a miss. It should have been in prod mid may, not MIA in late July. The brain drain at deep mind is a clear indicator that the people who know the most think that they can’t stay at the frontier.

Antigravity NEEDED to be game-changing. Without the stream of data that Claude, Codex, and Cursor enjoy there is little chance of getting an effective reinforcement learning loop. For the first time in its history, GOOG is at a meaningful data disadvantage, and apparently a cultural one as well.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#472

Pricing per million input/output tokens: 2.5 Flash: $0.3 / $2.5 3.0 Flash: $0.5 / $3 3.5 Flash: $1.5 / $9 3.6 Flash: $1.5 / $7.5 --- 2.5 Flash-Lite: $0.1 / $0.4 3.1 Flash-Lite: $0.25 / $1.5 3.5 Flash-Lite: $0.3 / $2.5

I wonder if this is a plateau towards the real pricing of AI, if you layer in gemini-2.0-flash at $0.10 / $0.70 then its a 15x price increase to 3.5/3.6 flash. But it hasnt gone up again which is interesting.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#473
Google really needs to get their product strategy together. The discontinued gemini-cli and introduced antigravity-cli which is a downgrade IMO and the sooner they can partner up with AWS and release the gemini models via Bedrock the easier corporate/business which has strict data protection rules can use their models and make them available for internal engineers.

It's a one thing to research and improve the model, but if they ignore the ease of access and multi-availability of their models in different ways they are going to fall behind again.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#474

Earlier quoted context omitted.

That's the only explanation that makes sense. If it was frontier but cost or compute were limiting factors, they'd release it at an obscene price for the bragging rights. Google doesn't care that much about alignment, and I don't think it's likely to be significantly different than 3.5 anyway. The only reason it would need to be soft-canceled is if it's terrible, and has to end up in a ditch like Llama 4 to avoid sha…

3.5 pro was clearly a miss. It should have been in prod mid may, not MIA in late July. The brain drain at deep mind is a clear indicator that the people who know the most think that they can’t stay at the frontier. Antigravity NEEDED to be game-changing. Without the stream of data that Claude, Codex, and Cursor enjoy there is little chance of getting an effective reinforcement learning loop. For the first time in its…

[deleted]

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#475
post #27

Earlier quoted context omitted.

Because AA Coding "Index" consists only of two benchmarks (Terminal-Bench v2.1, SciCode) and generally fails to be meaningfully representative of agentic coding capabilities.

AA coding index has been updated to use DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA.

When? It literally says on the page for Gemini 3.6 Flash "Artificial Analysis Coding Index represents the weighted average of coding benchmarks in the Artificial Analysis Intelligence Index (Terminal-Bench v2.1, SciCode)"

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#476
post #358

Earlier quoted context omitted.

In terms of open models, Gemma 4 beats the pants off everything else to the point that paying for APIs becomes hard to justify. Qwen has the meme-share for coding, but it feels much less well rounded. I have no doubt that Google have both the infrastructure and the expertise to curb stomp everyone else, should they resolve in earnest to do so. Lest we forget, "Attention is All You Need" came from Google.

Interesting, glad to hear. We have gemma4 at work, and I was considering localhosting qwen, but gemma4 is so far behind the Opus and Fable I have at home that I've decided to hold off for another model release.

"We have a company provided Toyota at work but it is so far behind the Ferrari I rent at home that I've decided to hold off for another model release."

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#477
post #449

Earlier quoted context omitted.

You’re absolutely right and it’s heartening to see. I maintain a client with ~every provider you can think of and llama.cpp and it was really tiring the last few days to see people laundering other stuff through Kimi and Qwen. They’re not even open yet, the hype was based on their own blog posts, no one’s actually running these locally, the Qwen Max’s have never been open, Kimi’s API was 1/2 the speed the benchmarks…

"You’re absolutely right and it’s heartening to see" Damnit, I usually don't jump to LLM speech patterns, but this opening had me thinking you were a bot. But after checking your profile, I think you pass as human. I wonder when will be the time, this does not work anymore for me. (Creation date is a strong hint, but abandoned accounts can be hijacked)

Hehe, cheers, it really is funny & odd habit I have (usually when I'm in "everyone is wrong!" mode, haven't bothered to argue that, and see someone else arguing it :p)

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#478
post #107

Earlier quoted context omitted.

It's also very possible that they know their big model underperforms chatgpt 5.6 and fable by too much, so they are focusing on what they can get wins in like speed instead.

Yes agreed - I wrote this up a while back https://martinalderson.com/posts/whats-going-on-with-gemini/ My view then was they are optimising the models for inference ability on their own hardware AND use cases, which is often speed and time to first token. They've somehow seemed to end up with terrible compute shortages, which again is surprising given how good Google is at infra deployments AND have their own hardwar…

They include an LLM response with every single Google search, whether it is warranted or not. That scale is, my guess, many orders of magnitude higher than what OpenAI and Anthropic serve. And for Google none of these are paid interactions since their LLMs do not (YET) insert ads into the responses.

So my guess is that Google will continue having compute shortages until the Gemini enshittification starts.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#479
post #117

Earlier quoted context omitted.

This is the feeling i get too. Cant produce quality, but can produce something that is super fast...so take the wins where they are.

For a coding LLM specifically, when is fast a good tradeoff for quality?

A coding agent driven by a large LLM can delegate smaller tasks to a faster model. For example searching through the codebase for references, examples, or established patterns. They are treated as tools and don't pollute the main agent's context.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#480

Earlier quoted context omitted.

That and/or the business case isn’t as clear when serving enormous models? You’re constantly stuck in a red queen’s race where your profitability window is increasingly measured in weeks because the Chinese are right behind you. For small models (which are probably distilled from their big ones) you can serve them economically all the time and not hemorrhage money.

[flagged]

> is the same exact reason the Kimi crowd is wrong now. And it's very obvious that they're wrong, but they have intense emotional blinders on.

On the very link on the top comment of this thread, which I repost here:

https://artificialanalysis.ai/models/gemini-3-6-flash

Kimi K3 is ahead of Fable 5 on several benchmarks.

So basically the angle went from "China cannot ever compete" to "China is six months behind" to "China is six weeks behind" to "China is six days behind but that's because they're distilling" and now you're saying "Yup sure, Kimi K3 is ahead on several benchmarks but you cannot host it yourself so this thing will go absolutely nowhere".

I mean: is it not a bit early to draw conclusions? It's been days since a chinese model is ahead of the very best / frontier US model on several benchmarks and you compare it to models who were clearly behind on everything.

Give it some time.

Post reply on HN