Live data from Hacker News

Gemini 3.8 Flash and 3.8 Flash Cyber

blog.google

151–160 of 697 posts

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#151

Google keeps flashing everyone where everyone is expecting to get PRO'bed.

We also had GLM-5.3 flash and Qwen 3.8 Flash Next, everyone's getting flashed and I think it's a good trend.

Almost suspect that the rate of improvement to post-training is so fast that small models have an advantage - it takes much more compute to train a bigger model, so the flash models are just running in circles (well, not exactly of course) around the larger models right now.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#152
post #113

I've been using Gemini 3.7 for my personal trip planning app. Across multiple benchmarks, it ranks higher on everything I tried: - Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the best at taking a cluster of places and working out a visiting order. - Photo ranking (which photo should be the hero). Gemini can tell whether a photo is of the thing or of the vie…

One thing in your comment surprised me: "when a thing opens and closes" Why you would rely on the model's weights to know opening hours, instead of having the model call a web search tool to verify it on the official site?

Maybe the model does some tool calling on its own to figure out the times?

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#153
"The knowledge cutoff date for Gemini 3.8 Flash is March 2026 – users can expect updated information for some domains while in others they may experience the model’s knowledge is limited to January 2025 (in line with the Gemini 3 Model Family)."

Kind of wild that they haven't (successfully) pretrained a base model since Jan-25.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#154
Everyone is censoring models now with anything remotely resembling cyber or bio. I already have problems with my research in mathematical epidemiology because of that - both Sol and Fable simply refuse. They keep pushing people towards Chinese models that can be decensored.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#156
post #131

Currently top at https://deepswe.datacurve.ai - beating Opus 5! https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium! Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.

sidenote, but wow sonnet 5 is shockingly bad on this benchmark.

sonnet 5 is bad by almost any metric.

anthropic really needs something to address the cheaper end of the market before they get left behind. Sonnet 5 sucks, and Haiku hasn't been updated in a year. meanwhile we've got gemini flash, luna, and GLM5.3 all delivering 90% of the performance for a small fraction of the cost. paying $25/mTok is going to start looking pretty silly soon.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#158
post #50

And yet again another failed launch from Google. I pay for their AI plus Google one package to get more cloud storage (have no interest in their AI bundle but you have to pay). and all I see in the Gemini app is 3.6-flash

Google has been doing staged roll outs on all their products since forever.

What stage of the roll out are we where I don’t even see 3.7-flash which was released 2-3 weeks ago?

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#160
3.7 flash was by far the best model for image recognition tasks according to my benchmarks. 3.8 flash didn't regress any candidates and improved some specificity (positive ID of common name vs species name of exotic fruit, correct identification of cast/replica of artifact and statue) but is still relatively weaker (26/30) on esoteric public figures (Korean beatboxers). I'm going to have to make my benchmark harder.
Post reply on HN