Live data from Hacker News

Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

blog.google

341–350 of 616 posts

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#341
post #4

It is both less intelligent and more expensive than GLM-5.2, while being closed weight.

Is that statement based on token price? More and more it seems that $/token hides as much as it reveals. Token efficiency, tokenizer differences, etc. I'm not saying that you are wrong, I am just saying it is becoming a bit more difficult making statements like this without a bit more research.

It's based on the Artificial Analysis "Intelligence Index vs. Cost per Intelligence Index Task" here:

https://artificialanalysis.ai/#intelligence-comparison-tabs

Differences in token "density" are accounted for by pricing per task

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#342

Google somehow managed to snatch defeat from the jaws of success with their AI products. They literally forced me and my company out of Antigravity by phasing out AI Ultra subscription without any proper product follow-up. Antigravity IDE cannot even have poweruser subscriptions now from Google Workspace an Gemini Enterprise Agent Platform cannot be attached to Antigravity IDE. Gemini Enterprise Agent Platform has an…

> Google somehow managed to snatch defeat from the jaws of success with their AI products. HN lives in a bubble. I have German/Italian/Polish clients virtually all use Gemini and NotebookLM. Talking insurance, banking, consulting, legal. The real world doesn't look at pointless benchmarks on writing react tailwind crap, they are already google suite users, get the tools, test them and adopt them, end of story. It's g…

I have friends and family who use Gemini, but entirely because their Pixel phones came with a year of it for free. No other reason, and they will most likely never pay for it.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#343
post #305

Earlier quoted context omitted.

How does your comparison work? It places Gemini 3.6 Flash Medium above GPT 5.6 Sol High and Fable 5 Medium, which makes me skeptical because that... would be making headlines that I'm not seeing right now.

I have created various questions/tests and put the models through the same tests. I record whether the answers are correct, and the generation stats (costs, latencies, tokens used, etc.). I have no idea why the Gemini models do so well. I have recently added new tests, whose sole purpose was to find some cases on which Gemini 3 Flash fails (I don't like cherry-picking models or tests, but I also find it strange Gemin…

Interesting, well it'd be interesting to check out some individual examples where Gemini beat the others.

Also, would be great if you could add GPT 5.6 Sol XHigh and Fable 5 High as well, just to see if at least those beat Gemini which is currently your #1.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#344
post #305

Earlier quoted context omitted.

How does your comparison work? It places Gemini 3.6 Flash Medium above GPT 5.6 Sol High and Fable 5 Medium, which makes me skeptical because that... would be making headlines that I'm not seeing right now.

I have created various questions/tests and put the models through the same tests. I record whether the answers are correct, and the generation stats (costs, latencies, tokens used, etc.). I have no idea why the Gemini models do so well. I have recently added new tests, whose sole purpose was to find some cases on which Gemini 3 Flash fails (I don't like cherry-picking models or tests, but I also find it strange Gemin…

Oh, and I've also added weights to different categories, so Coding and Tool usage categories influence the score more. This done both to better account for how most people are being used, and also to reduce Gemini's dominance in general/domain specific knowledge.

So yes, Gemini models are at the top, even if I actually (not proud of it) tried to make tests that actually favour other coding-focused models.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#345
post #283

I have a side business selling custom fingerprint jewelry and I use gemini nano banana to clean up customer submitted fingerprint images. This was a step I used to do by hand at 10 - 15 minutes per image and nano banana is the first model that is able to do the task (it is astonishingly good at it). I can't wait to see what the next nano banana can do, hopefully its released soon.

Are your customers clearly informed that you're sending their immutable fingerprints to an AI service?

Yes this is extremely unresponsible if so. Fingerprints are legally protected biometric data in most juristictions.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#347
post #145

Google somehow managed to snatch defeat from the jaws of success with their AI products. They literally forced me and my company out of Antigravity by phasing out AI Ultra subscription without any proper product follow-up. Antigravity IDE cannot even have poweruser subscriptions now from Google Workspace an Gemini Enterprise Agent Platform cannot be attached to Antigravity IDE. Gemini Enterprise Agent Platform has an…

> Google somehow managed to snatch defeat from the jaws of success This is still very early days. Who is "on top" has flipped back and forth many times already. The next frontier model release (from whomever) will change things again.

I guess? Maybe I'm alone here but I don't feel like Fable is particularly more useful than Sonnet most of the time. I feel like the LLMs are good enough for the majority of uses and the hyper expensive premium ones are way into diminishing returns. At this point with Kimi being as good as it is, if they jack up the price any more I'll just go open source.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#348

Earlier quoted context omitted.

As someone who's been using Workspace as a personal email account for over a decade this has been such a struggle forever. Just lots of odd limitations to feature sets all over the place. When they swapped Google Assistant for Gemini as the default voice provider in Android Auto it was so annoying. My wife's non-work space account can get Gemini to do the normal things like play music and what not, but my Workspace o…

I'm in the same situation. But I was shocked discovering it goes both ways: many new Gemini functionalities are only accessible using a consumer account instead of a Workspace account. Also, Gemini is now the only major AI assistant with no support for MCP connectors. Instead of adding this to the core product, like ChatGPT and Claude did, somebody at Google decided that it was smarter to add this fundamental feature…

Maybe because they want to train on your data? Workspace AIs are not trained on your corporate data.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#349
post #305

Earlier quoted context omitted.

I have created various questions/tests and put the models through the same tests. I record whether the answers are correct, and the generation stats (costs, latencies, tokens used, etc.). I have no idea why the Gemini models do so well. I have recently added new tests, whose sole purpose was to find some cases on which Gemini 3 Flash fails (I don't like cherry-picking models or tests, but I also find it strange Gemin…

Interesting, well it'd be interesting to check out some individual examples where Gemini beat the others. Also, would be great if you could add GPT 5.6 Sol XHigh and Fable 5 High as well, just to see if at least those beat Gemini which is currently your #1.

I don't like to divulge tests, but one of them is a chess puzzle.

> would be great if you could add GPT 5.6 Sol XHigh and Fable 5 High as well

I would like too, but I avoided them for several reasons:

1) Cost - this is a hobby project, those models would cost tens of dollars for each benchmark run, multiply this by tens or hundreds of models and ...

2) Time - the high models are already taking a really long answer to respond (5-10minutes per question). I run each question with 3 repeats (run the same test three times), so it would take 30 minutes per test. If I change my tests, methodology, or add a new test, it would take a really long time to run the benchmark. Also, I like having results immediately when a new model is released, now I can post within 30 minutes of a model's release the benchmark results.

3) High reasoning usually does WORSE on most tests - if you look at the leaderboard, it's sometimes counter-intuitive, but models with high or max reasoning usually do worse than medium and low. This is because the questions are quite targeted/direct, and the models overthink the question and miss the solution. Or the long thinking context makes them perform poorly. The generation tasks (SVGs/HTML animation) are usually better with longer reasoning, but short code fixes, trivia questions, puzzles, etc. are answered by low/med reasoning with more accuracy in general

Also, Fable is borderline un-testable, it refuses to answer many questions, so it scores poorly anyway.

Gemini scores 21/22 because it answers all tests, and it does them correctly, consistently. The only failed test is I think because it miscounted the lines in a file, when responding on which line the bug was in a code snippet.

Post reply on HN