Live data from Hacker News

Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

blog.google

561–570 of 616 posts

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#561
post #169
post #79

Earlier quoted context omitted.

I am growing tired of these pelicans posts every time a new model is published. Feels to me like low effort personal brand promotion. Just sharing my 2 cents.

You and a few other people, but enough people still appreciate the bit that I'm going to keep doing it. They're easy enough to skip - click the little "-" icon and you'll collapse the entire sub-thread.

Thank you for posting them (as well as all your other AI takes). It's appreciated here!

I feel like the human brain massively overweights negative feedback over positive, and that's even after accounting for the fact that internet discourse tends to mostly surface negative comments (whereas the enjoyers stay silent). By default I have to try hard not to take it personally whenever someone leaves a negative comment about my work.

So just doing my bit to say I appreciate your commentary on so much of the fast-evolving AI landscape. Helps me orient :)

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#563

Earlier quoted context omitted.

They include an LLM response with every single Google search, whether it is warranted or not. That scale is, my guess, many orders of magnitude higher than what OpenAI and Anthropic serve. And for Google none of these are paid interactions since their LLMs do not (YET) insert ads into the responses. So my guess is that Google will continue having compute shortages until the Gemini enshittification starts.

I don't think so. According to some very basic research there are around 8bn searches a day, or 250bn a month. Let's assume Google serves AI overviews on every SERP (they don't) and don't cache them (they do, afiak). And let's assume that each AI overview is 2000 tokens (blended input/output), that's 500T tokens a month. It's rumoured that anthropic is serving somewhere close to 10Q tokens a month. Now it may be that…

Well I asked the google AI mode thing what it thinks about your comment and it told me this (edited obviously):

"10 Quadrillion tokens a month means: 333 Trillion tokens per day and 3.85 Billion tokens generated/processed every single second, 24/7."

"At an incredibly cheap, subsidized infrastructure cost of $1 per million tokens, serving 10 Quadrillion tokens would cost Anthropic $10 Billion per month ($120 Billion a year) just in inference compute."

It also had this to say about how google's AI overview works: "Google doesn't just feed the LLM your 5-word search query. The system scrapes the top 10–20 web results, feeds thousands of words (tens of thousands of tokens of context) into the model, processes it, and then outputs the result."

Oh, and it does all of that in less than two seconds. Honestly, whatever Google is doing with its infrastructure is so far ahead of everyone else, I can't believe you fell for such an obvious lie.

There are also extremely obvious holes in your comment:

>Let's assume Google serves AI overviews on every SERP (they don't) and don't cache them (they do, afiak).

Try it out for yourself. Add a few random letters or punctuation. They cache nothing.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#564
post #375

Earlier quoted context omitted.

I choose fourth option. 4) googles big model just performs worse than K3 and GLM so they choose not to embarass themself. Like I love Gemini and use it a lot to one-shot whole MR with huge contexts, but its just much worse when its come to tool use and agentic coding.

or they just don't want compete in coding space ??? they have search,youtube,android,office suite like gmail,maps,spreadsheet etc coding is the least of their problem/priority

Considering how cheap the subscriptions are, it looks like agentic coding is a low margin business. If they can sell you a subscription for the chat, it is profitable, but if you try to use the subscription to its limits, you're probably making them lose money.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#565

Earlier quoted context omitted.

or they just don't want compete in coding space ??? they have search,youtube,android,office suite like gmail,maps,spreadsheet etc coding is the least of their problem/priority

Considering how cheap the subscriptions are, it looks like agentic coding is a low margin business. If they can sell you a subscription for the chat, it is profitable, but if you try to use the subscription to its limits, you're probably making them lose money.

absolutely, knowing OpenAI try to break into ads market tell the whole direction that pure AI is not that profitable tbh (especially with how cheap chinnese model are)

the integration on ecosystem is the bread are

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#566
post #189

Earlier quoted context omitted.

Here, my comparison of 3.6 Flash vs Sol vs Luna vs Terra: https://aibenchy.com/compare/google-gemini-3-6-flash-medium/...

You should really provide more on your methodology because as it stands, it really doesn't pass the sniff test. GPT-5.6 Sol on Low beats Fable Medium by 10% and Gemini-3.6 Flash then beats them both? Fable is number 20? This does not match any lived experience or developer experience. It'd be helpful to know _what_ you're testing and break that out by dimension. You mention randomly selected questions. How does that…

Thanks for the feedback, really good points!

There is some short info about the methodology here: https://aibenchy.com/methodology/

> GPT-5.6 Sol on Low beats Fable Medium by 10% > Fable is number 20

Fable loses a lot of points because it often refuses to answer questions. Asking a basic tool-usage challenge, Fable responded with refusal: "This request triggered restrictions on violative cyber content and was blocked under Anthropic's Usage Policy. To learn more, see https://platform.claude.com/docs/en/build-with-claude/refusa...." Even in practice, you ask Fable something trivial, and it refuses to respond. I think the score accurately represents how the model is behaving in real-world usage.

> Gemini-3.6 Flash then beats them both Gemini models are the most intelligent overall. The tasks are not coding-only. Gemini excels in general knowledge and domain specific knowledge. Gemini models, even old ones, still top many charts on specific use-cases[0][1]. Depending on how you weigh those cases, the leaderboard order can vary quite drastically, as some models are very strong in some domains and weak in others.

> You mention randomly selected questions. How does that work? Randomly selected, means I have manually created the questions/challenges to span across various domains and agentic surfaces. Questions vary from coding tasks, tool usage, trivia questions, chess puzzles, car-wash-like challenges and more.

> With n=22 and binary pass/fail Each test is run 3 times, so in total we have 66 tasks. Also, apart from correct/wrong answer, the final score also includes the pass rate for each test (how many out of three attempts), how good the reasoning is (they have a hidden reasoning score where available) and other small factors. Also, some tests in some categories involve a series of tasks/requirements (i.e. implement this function, call it, do some processing on the result, combine the result with some built-in knowledge data, etc.).

I do agree that 22 tests isn't that much, and I'm slowly adding more, but even without the leaderboard part, the comparison feature is what's I think is most useful. You can see for the exact same tasks, which models do better, which do it faster, which cost less, etc.

> _what_ you're testing and break that out by dimension There is a category breakdown for the test results, so you can see and which sort of tasks models fail.

Everything aside, when you manually ask a model to test its capabilities, I don't think it takes many questions to realise how good/bad that model is. Sometimes one prompt is enough, you ask it to do something, and see how it reasons about it, how fast it does it, how efficient the steps are and how good the result is. Yes, the performance may vary across tasks, but I'm pretty sure if you did a blind test with a chatbot, you could easily realise how good the model is in just a few questions/tasks.

I think no benchmark is perfect, mine is far from it, but it's simply another different, independent data-point. Apart from that, I made this for myself, and I'm using it myself. I don't trust that all popular benchmarks are not in the training data, and I think many benchmark the wrong things which don't correlate to how I use the models day-to-day myself. I just made the results publicly available, in case any one else benefits from it. I've probably spent thousands in LLM costs, and probably more than 100 hour building this, without benefiting in any way from it (outside of the joy of building it and me using it personally to compare models); as long as models cost stays reasonable, I'll keep building it and test new models as soon as they are released.

[0]: https://x.com/browser_use/status/2079602472516264010/photo/1 [1]: https://artificialanalysis.ai/evaluations/mmmu-pro#mmmu-pro-...

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#567

Earlier quoted context omitted.

That and/or the business case isn’t as clear when serving enormous models? You’re constantly stuck in a red queen’s race where your profitability window is increasingly measured in weeks because the Chinese are right behind you. For small models (which are probably distilled from their big ones) you can serve them economically all the time and not hemorrhage money.

[flagged]

Honestly go use Deepseek v4 Flash, the data hosting in PRC is it's only downfall. It truly is an excellent model at the moment. It is opensource thankfully but there is friction there and this won't be as popular.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#568

Earlier quoted context omitted.

It’s really surprising. When Apple announced the multi-billion dollar deal with Google to power Apple Intelligence I thought great things were coming. Instead we are getting more and more bad news: delayed Pro models and AI leadership leaving. I wonder if Apple know something the rest of us don’t know or if they are already regretting their decision.

Besides Apple apparently making Siri AI model agnostic, the choice to go with Google was almost certainly for practical reasons. Google is a low-risk established player that already has a long work history with Apple. Google also isn't in an existential battle to establish themselves, Gemini still amounts to just another project at Google. There is tangible non-zero risk that either OAI or Anthropic will be gone in 5…

They also famously hate nvidia since 2008 and would prefer to use TPUs, at least historically

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#569

Earlier quoted context omitted.

[flagged]

Why was this flagged? This is absolutely true, Americans seem to not actually want to use Chinese models at least for coding, maybe for other inference use cases but I haven't seen it. No one I know uses anything but OpenAI and Anthropic even if Chinese models are better or cheaper in many use cases.

First 6 of the most used models on openrouter currently are Chinese; that's true for code generation and other coding-relevant tasks too when ranked by share of tokens.

https://openrouter.ai/rankings#top-models

Of course openrouter is not representative because most users directly go to the model provider but it still proves your claim is very far from "absolutely true".

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#570
post #75

A couple tidbits: > Beyond today’s releases, Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it’s ready. > We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress.

I guess they meant to release Gemini 3.5 Pro shortly after 3.5 Flash, but then Mythos/Fable and later GPT-5.6 came out with higher performance than 3.5 Pro, so the managers decided not to release it.

That reasoning didn't stop them from releasing this batch of models, admittedly that may be less face to lose away from the flagship position.
Post reply on HN