Live data from Hacker News

Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

blog.google

541–550 of 616 posts

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#541
post #4

It is both less intelligent and more expensive than GLM-5.2, while being closed weight.

It'd be interesting to know how much the Intelligence as a Service angle serves as a value-add in the minds of Google's executives. You can get decent open-weight models now. That's not difficult. The difficulty is 1) running them and 2) compliance. My company runs Claude on GCP's Vertex AI solution. We're in the US healthcare IT space, so the models need to be from somewhere that American healthcare agencies and com…

[flagged]

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#544

I wonder how big the Pro model is that Google is using behind the scenes to train these smaller ones. Going on baseless speculation, the lack of accompanying pro models with these flash releases either means: 1) the model is too big to be economical, 2) google doesn't have the compute to serve the big model, 3) their big model has too many alignment issues to serve to the public. edit: looks like benchmarks are up on…

I suspect the entire 3.x family is fundamentally problematic and we'll need to see an architecture change before they're half decent like back in the 2.5 days again.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#545
post #189
post #13

It's a bit disheartening to see no comparison to other models here - and I'm not sure this pushes the curve anywhere. 3.6 flash is more expensive than GLM 5.2 - but seemingly worse, although this post is really light (lite?) on details. It seemed for a time that Google had finally gotten the ball rolling, but I'm doubting that more and more as time passes. We'll see what happens with 3.5 pro I suppose.

Here, my comparison of 3.6 Flash vs Sol vs Luna vs Terra: https://aibenchy.com/compare/google-gemini-3-6-flash-medium/...

You should really provide more on your methodology because as it stands, it really doesn't pass the sniff test. GPT-5.6 Sol on Low beats Fable Medium by 10% and Gemini-3.6 Flash then beats them both? Fable is number 20?

This does not match any lived experience or developer experience.

It'd be helpful to know _what_ you're testing and break that out by dimension. You mention randomly selected questions. How does that work?

With n=22 and binary pass/fail, the 95% confidence interval on a pass rate spans roughly (+-)15-20 percentage points. There's just not enough data ironically, for this leaderboard to mean anything. Ranks #5 through #25 are statistically indistinguishable

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#546
post #493

Earlier quoted context omitted.

"As of October [2025], OpenAI's compute margins reached 70%, up from 52% at the end of 2024 and double the rate in January 2024, [The Information] said, citing a person familiar with the figures." https://www.bloomberg.com/news/articles/2025-12-21/openai-se... As for Anthropic, the rumors I remember seeing for their API margins were more like 85-90%, but I don't have a reference at hand for those. But once you know t…

It says the original report was in the Information, which I can't see, but I'm skeptical that they includes the training cost? And how much that changes the figure?

That's the profit margin on inference, not overall. Each model does end up being profitable over its lifetime, but the money they're making is being immediately churned into buying more data centers & the training for the next giant model up, so they're not profitable overall at the moment. It's a bit like how Amazon kept churning their profits into more growth instead of taking the profit early.

That said, Anthropic has supposedly crossed over into profitability and made $1 Billion in profit so far this year, in the lead up to their IPO. Being profitable sounds good for launching on the stock market! But as a customer, that's noticeable in the downtime due to lack of compute, and only getting 50% access to Fable.

OpenAI might not be profitable, but they've got so much compute access that they've been able to give their customers full access to Sol, and as a result they've almost doubled their Codex subscriber base in the last two weeks (6 million on July 12, 10 million on July 21 - that would be an extra $1-$10 Billion in Annual Recurring Revenue that they've gained in just these 2 weeks). Doing the unprofitable thing in the short term can result in outsized rewards in the long term.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#547

Earlier quoted context omitted.

Aren’t the subscriptions extremely subsidized and burning cash for Anthropic and OpenAI? A reasonable explanation is they’re simply abstaining from the war of attrition, especially given cheaper comparable models are breaking the illusion that the “frontier of intelligence” has any kind of per token margin.

This is hotly debated and completely unclear. Let's say Anthropics Opus models cost the same to serve as GLM 5.2. GLM 5.2 is 4.4$/MTok while Opus is 5.6 times more expensive. Assume that GLM 5.2 is served at essentially zero margin. Then Anthropic has >80% margin on API pricing. So even if an average person with a subscription pays only 20% of the API price of their usage, Anthropic makes money on subscriptions. And…

I know many, many engineers who are paying something like 2-5% (via subscription) of what their usage would cost if billed by API tokens.

I know some down to about 1% ($200 Max plan vs $20k in tokens per month)

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#548

I wonder how big the Pro model is that Google is using behind the scenes to train these smaller ones. Going on baseless speculation, the lack of accompanying pro models with these flash releases either means: 1) the model is too big to be economical, 2) google doesn't have the compute to serve the big model, 3) their big model has too many alignment issues to serve to the public. edit: looks like benchmarks are up on…

I wish someone would convince Google to may be leave the Google search be without AI responses and use all their resources for a Gemini subscription/API..

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#549
post #79
post #71

Pelicans for 3.6 Flash and 3.5 Flash-Lite (Cyber isn't available to me through the API yet.) https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

I am growing tired of these pelicans posts every time a new model is published. Feels to me like low effort personal brand promotion. Just sharing my 2 cents.

[deleted]

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#550

Earlier quoted context omitted.

This is hotly debated and completely unclear. Let's say Anthropics Opus models cost the same to serve as GLM 5.2. GLM 5.2 is 4.4$/MTok while Opus is 5.6 times more expensive. Assume that GLM 5.2 is served at essentially zero margin. Then Anthropic has >80% margin on API pricing. So even if an average person with a subscription pays only 20% of the API price of their usage, Anthropic makes money on subscriptions. And…

I know many, many engineers who are paying something like 2-5% (via subscription) of what their usage would cost if billed by API tokens. I know some down to about 1% ($200 Max plan vs $20k in tokens per month)

And I know many people that don't. That have a 20 or 100 dollar subscription for very bursty workflows with months where they barely use tokens.

Not every subscriber is a full time SWE. In fact most professional SWEs will be on enterprise plans and thus not get subscriptions at all.

I think it's very plausible that subscriptions are overall losing money. But we simply don't know.

Post reply on HN