Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
581–590 of 616 posts
Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
#582I wonder how big the Pro model is that Google is using behind the scenes to train these smaller ones. Going on baseless speculation, the lack of accompanying pro models with these flash releases either means: 1) the model is too big to be economical, 2) google doesn't have the compute to serve the big model, 3) their big model has too many alignment issues to serve to the public. edit: looks like benchmarks are up on…
It used to be that you could find the edges of the training set pretty easily. No longer.
Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
#583Earlier quoted context omitted.
It's rumored that Gemini 3.5 flash has a >50% margin, and I'd imagine 3.6 flash is even higher. I do not think OpenAI or Anthropic are actively chasing margins - though, Anthropic is supposed to be profitable on some form of non-GAAP accounting... I suspect Google isn't really interested in seeing how far it can get dragged into a race of selling dollars for $0.25, and is more interested to see if it can stay in the…
It kind of doesn't make sense though, because typically a large org like Google can afford to crush competitors on pricing. They could probably even go toe to toe with chinese model pricing for years without feeling it. Maybe they don't want to price war with the other labs so they can comfortably maintain healthy margins on selling them compute?
Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
#584Pelicans for 3.6 Flash and 3.5 Flash-Lite (Cyber isn't available to me through the API yet.) https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
#585Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
#586Google seems to have anorexia when it comes to model intelligence. They have an internal hard constraint on price per token it seems, and they are trying to squeeze out intelligence with limited compute. I wonder if there is something with their TPU cycles that makes them want to postpone training a new model. My guess is that they have been on the same base model for 6 months and they may have waited for the next ge…
Could it be that they have to serve their models to billions of users?
Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
#587I wonder how big the Pro model is that Google is using behind the scenes to train these smaller ones. Going on baseless speculation, the lack of accompanying pro models with these flash releases either means: 1) the model is too big to be economical, 2) google doesn't have the compute to serve the big model, 3) their big model has too many alignment issues to serve to the public. edit: looks like benchmarks are up on…
They are likely deliberately avoiding the SoTA race for a few reasons:
1. Their best models are marginally better than current SoTA releases. 2. They'd like to let Ant/OAI make mistakes with safeguards / let them get the regulatory heat. The unknown unknowns are huge with SoTA models (eg OAI accidentally hacking huggingface) and they are protecting their reputation. 3. They want to encourage companies to become cost conscious because they can likely win on price in the long run. Getting market share in "quantity beats quality" workflows forces companies to establish processes to choose the "cheapest acceptable model", which is a good environment for Google.
Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
#588Earlier quoted context omitted.
While them fixing token bloat on 3.5 Flash is good work. That paragraph was the real highlight. Hopefully 3.5 Pro is soon, and that Gemini 4 can be here end of year and finally have an updated knowledge cutoff.
What is a “knowledge cutoff” ? —Ah, got it, it knows more about recent times.
Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
#589Earlier quoted context omitted.
It’s really surprising. When Apple announced the multi-billion dollar deal with Google to power Apple Intelligence I thought great things were coming. Instead we are getting more and more bad news: delayed Pro models and AI leadership leaving. I wonder if Apple know something the rest of us don’t know or if they are already regretting their decision.
Besides Apple apparently making Siri AI model agnostic, the choice to go with Google was almost certainly for practical reasons. Google is a low-risk established player that already has a long work history with Apple. Google also isn't in an existential battle to establish themselves, Gemini still amounts to just another project at Google. There is tangible non-zero risk that either OAI or Anthropic will be gone in 5…
Ridiculous statement, pretty much all LLM tooling is model agnostic.
Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
#590Earlier quoted context omitted.
I know many, many engineers who are paying something like 2-5% (via subscription) of what their usage would cost if billed by API tokens. I know some down to about 1% ($200 Max plan vs $20k in tokens per month)
And I know many people that don't. That have a 20 or 100 dollar subscription for very bursty workflows with months where they barely use tokens. Not every subscriber is a full time SWE. In fact most professional SWEs will be on enterprise plans and thus not get subscriptions at all. I think it's very plausible that subscriptions are overall losing money. But we simply don't know.
(and if you are using per-token billing as an enterprise user without first maxing out a premium seat you are very silly)