Live data from Hacker News

Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

blog.google

581–590 of 616 posts

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#582

I wonder how big the Pro model is that Google is using behind the scenes to train these smaller ones. Going on baseless speculation, the lack of accompanying pro models with these flash releases either means: 1) the model is too big to be economical, 2) google doesn't have the compute to serve the big model, 3) their big model has too many alignment issues to serve to the public. edit: looks like benchmarks are up on…

I've been using 3.5 flash in Android Studio on a Dart/Flutter project, some of it pretty complex and using brand new API's and native code for agentic tool calling in the app.

It used to be that you could find the edges of the training set pretty easily. No longer.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#583

Earlier quoted context omitted.

It's rumored that Gemini 3.5 flash has a >50% margin, and I'd imagine 3.6 flash is even higher. I do not think OpenAI or Anthropic are actively chasing margins - though, Anthropic is supposed to be profitable on some form of non-GAAP accounting... I suspect Google isn't really interested in seeing how far it can get dragged into a race of selling dollars for $0.25, and is more interested to see if it can stay in the…

It kind of doesn't make sense though, because typically a large org like Google can afford to crush competitors on pricing. They could probably even go toe to toe with chinese model pricing for years without feeling it. Maybe they don't want to price war with the other labs so they can comfortably maintain healthy margins on selling them compute?

Why would they go into price war when they are compute constrained?

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#584
post #71

Pelicans for 3.6 Flash and 3.5 Flash-Lite (Cyber isn't available to me through the API yet.) https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

I only see a hat under 3.5 Flash-Lite. Is it a model issue or a rendering issue?

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#585
post #468

Earlier quoted context omitted.

the person you were replying to says absolutely nothing about using it to write code.

https://news.ycombinator.com/item?id=48766580

brother thats not the comment you replied to

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#586

Google seems to have anorexia when it comes to model intelligence. They have an internal hard constraint on price per token it seems, and they are trying to squeeze out intelligence with limited compute. I wonder if there is something with their TPU cycles that makes them want to postpone training a new model. My guess is that they have been on the same base model for 6 months and they may have waited for the next ge…

Could it be that they have to serve their models to billions of users?

It's interesting to see a competitive landscape shift out from under the early adopters. Anthropic appears to have anticipated this, but everyone else is going to get caught out by the major platform providers' multitude of distribution channels and use cases. Meta might also be an exceptional case if they can find a way to goose up ad effectiveness.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#587

I wonder how big the Pro model is that Google is using behind the scenes to train these smaller ones. Going on baseless speculation, the lack of accompanying pro models with these flash releases either means: 1) the model is too big to be economical, 2) google doesn't have the compute to serve the big model, 3) their big model has too many alignment issues to serve to the public. edit: looks like benchmarks are up on…

What I think is really going on is an attempt to segment the market in favor of Google's strengths. It's a bet that models are "good enough" for many use cases even before they reach human-level intelligence, and Google is trying to capture workflows where quantity beats quality.

They are likely deliberately avoiding the SoTA race for a few reasons:

1. Their best models are marginally better than current SoTA releases. 2. They'd like to let Ant/OAI make mistakes with safeguards / let them get the regulatory heat. The unknown unknowns are huge with SoTA models (eg OAI accidentally hacking huggingface) and they are protecting their reputation. 3. They want to encourage companies to become cost conscious because they can likely win on price in the long run. Getting market share in "quantity beats quality" workflows forces companies to establish processes to choose the "cheapest acceptable model", which is a good environment for Google.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#588
post #572

Earlier quoted context omitted.

While them fixing token bloat on 3.5 Flash is good work. That paragraph was the real highlight. Hopefully 3.5 Pro is soon, and that Gemini 4 can be here end of year and finally have an updated knowledge cutoff.

What is a “knowledge cutoff” ? —Ah, got it, it knows more about recent times.

I’d just like to add, and this may interest you both, that I’m not disagreeing with either of you. When 3.5 Flash first showed up in AI Studio, its model card said it had a March 2026 knowledge cutoff. After the backlash over it apparently knowing nothing past December 2024, the card was changed to read: “Knowledge cutoff: Unknown.” Maybe the timing was coincidental, but I doubt it.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#589

Earlier quoted context omitted.

It’s really surprising. When Apple announced the multi-billion dollar deal with Google to power Apple Intelligence I thought great things were coming. Instead we are getting more and more bad news: delayed Pro models and AI leadership leaving. I wonder if Apple know something the rest of us don’t know or if they are already regretting their decision.

Besides Apple apparently making Siri AI model agnostic, the choice to go with Google was almost certainly for practical reasons. Google is a low-risk established player that already has a long work history with Apple. Google also isn't in an existential battle to establish themselves, Gemini still amounts to just another project at Google. There is tangible non-zero risk that either OAI or Anthropic will be gone in 5…

> Besides Apple apparently making Siri AI model agnostic

Ridiculous statement, pretty much all LLM tooling is model agnostic.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#590

Earlier quoted context omitted.

I know many, many engineers who are paying something like 2-5% (via subscription) of what their usage would cost if billed by API tokens. I know some down to about 1% ($200 Max plan vs $20k in tokens per month)

And I know many people that don't. That have a 20 or 100 dollar subscription for very bursty workflows with months where they barely use tokens. Not every subscriber is a full time SWE. In fact most professional SWEs will be on enterprise plans and thus not get subscriptions at all. I think it's very plausible that subscriptions are overall losing money. But we simply don't know.

Claude Code has enterprise subscriptions. No $200 max plan but you can buy $100 premium seats.

(and if you are using per-token billing as an enterprise user without first maxing out a premium seat you are very silly)

Post reply on HN