Live data from Hacker News

Gemini 3.8 Flash and 3.8 Flash Cyber

blog.google

101–110 of 699 posts

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#101
Looks like the strategy of regular updates with incremental improvements is working out well. Interestingly, the biggest jump in Artificial Analysis Intelligence Index score is for reasoning level Medium ( 3.7 was 51, 53, 57 for Low, Medium and High, 3.8 is 52,57, 59 respectively). I think scores at lower reasoning levels are more indicative of model capability since higher reasoning levels are focussed on benchmaxxing. We use the lowest reasoning level in production with good results.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#102
post #18

Wait, I didn't realize 3.7 Flash was already beating Sol on a bunch of the benchmarks. Isn't it a way smaller models?

They're quite selective in benchmarks, c.f. notably only bad one is 10% on TerminalBench. It's a really addled model, one time I said "Hi" and it built out a 4 panel hello world app with (fake) weather, a todo list, and a couple other things I forgot. I wouldn't be comfortable saying "ignore the #s!" except when I complained it was trash and way overcooked on agentic coding yet not good at it, and a couple DeepMind ML people liked the tweet.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#103

Currently top at https://deepswe.datacurve.ai - beating Opus 5! https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium! Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.

Wait a week with your judgement - most likely, Google is just bench-maxing very hard. If you look at the previous Flash models and the announcement on Google I/O, it was an absolute disaster. Reality diverged very much from the marketing (supposedly great benchmarks).

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#106

Earlier quoted context omitted.

If by reckless you mean commit, push, deploy without me asking it to, the I agree!

Respectfully: If it's able to deploy without you asking it to, that's a you problem. There are no safeguards?

I told it “don’t betray me” in my prompt and it still stabbed me in the back.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#107

Earlier quoted context omitted.

If by reckless you mean commit, push, deploy without me asking it to, the I agree!

Respectfully: If it's able to deploy without you asking it to, that's a you problem. There are no safeguards?

That's exactly how you get 'you are right, I deleted the production DB to apply the new schema when I should have written a migration'

That said, I do trust Opus and Fable enough to let them deploy to staging. Great for debugging. Just don't give them keys for prod

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#108

Currently top at https://deepswe.datacurve.ai - beating Opus 5! https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium! Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.

On artificial analysis it's only equal to opus 5 medium effort. Opus 5 max scores 63. Further, opus 5 medium outputs 4x fewer tokens to achieve the same result, negating a lot of the speed difference.

A comparison to an artificial score and a comparison to “the same task”

These folks must laugh themselves to sleep. This whole industry hoodwinked the masses. It’s impressive.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#109

One place where I find the Flash models surprisingly bad is Google Search's "AI Mode". A recent example - I searched for how to unsubscribe from Pearson emails. Google Search "AI Mode" confidently gave me a sequence of steps along the lines of Settings > Profile > Email preferences > Unsubscribe. Of course, I looked for an unsubscribe link before asking Google. None of those options existed. The correct answer was th…

It's not the models, it's the guardrails.

It's obvious that the Google Search AI Mode encourages the model to give an answer without spending unnecessary cycles investigating deeply.

They also heavily encourage keeping the context short. For example, it will remove the option to start a new turn after a small number of turns, depending on the topic.

It definitely makes things up all the time, but it gets it right surprisingly often. I really like it.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#110
post #64

Pelicans (thinking effort high, medium, low): https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - high cost 8.9742 cents Here are the 3.7 pelicans for comparison: https://tools.simonwillison.net/markdown-svg-renderer.html?u... - high cost 8.4387 cents (I think thinking level low is a regression on 3.8 compared to 3.7.)

I mean no offense but these pelicans are a bit tiresome and a very meaningless benchmark. There's no real difference between any of these svgs across models and model versions anymore.

It is more fun than serious at this point. Don't overthink it :)
Post reply on HN