Live data from Hacker News

Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

blog.google

521–530 of 616 posts

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#521

Man I love Gemini models but these kind of pricing increase is just insanse. I have a little product and I have to keep increasing the price and reduce the limits because of this non-sense, and they did not even let us use the old models in near future, so I forced to update to the new model with basically no to little improvement because I don't even need that much. Google if you can read this, it okay to release ne…

Do not base products on models that are not open-weights. Doing it is like building a product on someone else's platform, you are entirely at their mercy, and even when they don't have any reason to hurt you, you are tiny enough that if any policy they want to enact hurts you as a side effect, no-one is going to care.

You don't have to self-host the open-weights model, you just need to be able to source it from multiple providers.

Using the closed vendor models maybe made sense when open-weight models lagged so far behind, but that time is now gone.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#522
post #117

Earlier quoted context omitted.

This is the feeling i get too. Cant produce quality, but can produce something that is super fast...so take the wins where they are.

We don't have enough fast models, so I see this as a positive. I just test drove Gemini Flash Lite and it's crazy fast.

we need fast + cheap AI model, those gemini 2.0 flash is superb

like it literally pennies

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#523
post #375

I wonder how big the Pro model is that Google is using behind the scenes to train these smaller ones. Going on baseless speculation, the lack of accompanying pro models with these flash releases either means: 1) the model is too big to be economical, 2) google doesn't have the compute to serve the big model, 3) their big model has too many alignment issues to serve to the public. edit: looks like benchmarks are up on…

I choose fourth option. 4) googles big model just performs worse than K3 and GLM so they choose not to embarass themself. Like I love Gemini and use it a lot to one-shot whole MR with huge contexts, but its just much worse when its come to tool use and agentic coding.

or they just don't want compete in coding space ???

they have search,youtube,android,office suite like gmail,maps,spreadsheet etc

coding is the least of their problem/priority

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#524
post #454

Earlier quoted context omitted.

Knowledge Graph + automated transcriptions of almost every YouTube video = giant untapped moat of data

It's not exactly untapped, my AI company has scraped YouTube transcripts for 3 years now for RAG

Yeah and I've used this site for years when trying to recall a lecture I watched but only remember bits of

https://filmot.com/

(it lets you search youtube by transcripts)

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#525
post #107

Earlier quoted context omitted.

It's also very possible that they know their big model underperforms chatgpt 5.6 and fable by too much, so they are focusing on what they can get wins in like speed instead.

That's the only explanation that makes sense. If it was frontier but cost or compute were limiting factors, they'd release it at an obscene price for the bragging rights. Google doesn't care that much about alignment, and I don't think it's likely to be significantly different than 3.5 anyway. The only reason it would need to be soft-canceled is if it's terrible, and has to end up in a ditch like Llama 4 to avoid sha…

It's also what Pichai literally said in a recent interview, that Google is not doing well in coding and agentic tasks.

https://www.searchenginejournal.com/pichai-says-google-is-a-... (link to the actual podcast interview source within, this has a summary)

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#526

Earlier quoted context omitted.

Yeah that’s the joke :p

I've never questioned the word lite before because it's existed my whole life.. So does it make sense? Why does it exist and where does it come from? More than coming from "light".

There's an Merriam-Webster article on it: https://www.merriam-webster.com/wordplay/lite-word-history

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#527

Earlier quoted context omitted.

That and/or the business case isn’t as clear when serving enormous models? You’re constantly stuck in a red queen’s race where your profitability window is increasingly measured in weeks because the Chinese are right behind you. For small models (which are probably distilled from their big ones) you can serve them economically all the time and not hemorrhage money.

[flagged]

Why was this flagged? This is absolutely true, Americans seem to not actually want to use Chinese models at least for coding, maybe for other inference use cases but I haven't seen it. No one I know uses anything but OpenAI and Anthropic even if Chinese models are better or cheaper in many use cases.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#528
post #223

Earlier quoted context omitted.

I don't think it's that early tbh, agentic coding has ~90% adoption in the US. Claude Code has largely won individual developer mindshare and has been on top ever since it came out. The benchmarks change, but almost nobody opts to use anything other than Claude IME when I ask them. Enterprise is more competitive since they care about costs and other things, but developers leaning towards Claude puts a thumb on the sc…

> Claude Code has largely won individual developer mindshare and has been on top ever since it came out. Claude Code's success is not due to the agent but because the model is considered the best for programming and is very heavily subsidized, compared to pay as you go API prices. Consumers and Enterprise are not really locked in and will go where it makes the most sense. I think they have almost no loyalty by actual…

Nah, these harnesses like Claude Code and Codex are very sticky, I don't see many coworkers switching often, as they have set workflows for their particular harness. It's like vim and emacs, once you pick one it's unlikely you'd switch especially as the models are all "good enough" by now.

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#529

My hunch is Google is trying to integrate a fast and relatively cheap AI across search and every other surface of their product suite. And for that objective, a model that can move faster while being accurate and cheap enough is more important to them than producing a frontier class heavyweight model.

This. Give me cheap tokens that produce accurate information and the deal is done

I've found that having good source material (markdown, dependency source, search results for agents) for the models to draw on significantly improves information accuracy. Definitely worth investing in this side of "harness engineering", don't rely on facts burned into weights

Re: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

#530

Earlier quoted context omitted.

This is hotly debated and completely unclear. Let's say Anthropics Opus models cost the same to serve as GLM 5.2. GLM 5.2 is 4.4$/MTok while Opus is 5.6 times more expensive. Assume that GLM 5.2 is served at essentially zero margin. Then Anthropic has >80% margin on API pricing. So even if an average person with a subscription pays only 20% of the API price of their usage, Anthropic makes money on subscriptions. And…

Sure, but then why wouldn’t I use GLM 5.2 at cost or K3? I guess that’s the big question, will people pay a big margin long term to use their end products / models or will AI tokens be commoditized by many competing players. For coding if I had to pay API costs I’d switch in a heartbeat, enterprise maybe more reluctant?

Enterprise contracts with American companies over foreign ones. Enterprise is where OpenAI and Anthropic make most of the money.
Post reply on HN