Live data from Hacker News

Gemini 3.8 Flash and 3.8 Flash Cyber

blog.google

91–100 of 699 posts

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#92

Currently top at https://deepswe.datacurve.ai - beating Opus 5! https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium! Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.

A fifth of the cost of Opus 5! Google is certainly pushing the completion with this.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#93

One place where I find the Flash models surprisingly bad is Google Search's "AI Mode". A recent example - I searched for how to unsubscribe from Pearson emails. Google Search "AI Mode" confidently gave me a sequence of steps along the lines of Settings > Profile > Email preferences > Unsubscribe. Of course, I looked for an unsubscribe link before asking Google. None of those options existed. The correct answer was th…

I think that's just a limitation on the size of the model. I'm pretty sure that they use a pretty small model in those summaries to save money, which naturally makes them a little less smart.

[dead]

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#94
post #56

Earlier quoted context omitted.

The benchmark also doesn't include speed. You almost think something has gone wrong when using it because it returns full responses so incredibly fast.

Not just speed, also reliability. IME, Gemini's speed and quality doesn't degrade badly during weekday working hours compared to OAI, and especially Anthropic.

I've had Gemini model API use degrade the most out of OAI/Anthropic/Google (often "over capacity" vs true failures)

Not sure on consumer/product use though

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#95

Currently top at https://deepswe.datacurve.ai - beating Opus 5! https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium! Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.

We'll see about that. I suspect benchmaxxing as all the labs do as I haven't found Gemini models to be nearly as good in agentic engineering compared to Claude or GPT models.

If anything, gemini models are the least benchmaxxed out of any lab, IMO.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#96
post #92

Currently top at https://deepswe.datacurve.ai - beating Opus 5! https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium! Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.

A fifth of the cost of Opus 5! Google is certainly pushing the completion with this.

Gemini hasn't failed me for personal usage yet. I haven't had the opportunity to use it at work.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#97
The most interesting thing about the Gemini models is still their multi-modal support: they accept audio and video input, OpenAI and Anthropic's flagships are still image-only.

Gemini Flash is also pretty cheap, so it's a great family for performing media analysis, like extracting structured data from images and video.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#98
post #79

One place where I find the Flash models surprisingly bad is Google Search's "AI Mode". A recent example - I searched for how to unsubscribe from Pearson emails. Google Search "AI Mode" confidently gave me a sequence of steps along the lines of Settings > Profile > Email preferences > Unsubscribe. Of course, I looked for an unsubscribe link before asking Google. None of those options existed. The correct answer was th…

https://www.pearson.com/privacy-center/privacy-notices/full-... >We will not send marketing emails to a user who has opted out of receiving them. Any marketing communications we send will include an unsubscribe link at the end of the email. I don't think this is AI's fault. This is Pearson's publishing incorrect information and the only way to really know they are a bunch of lying assholes is to have an account and t…

[dead]

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#100

Earlier quoted context omitted.

Latest rumor is that 3.5 pro was struggling to be meaningfully better than flash, since iterations on flash were moving much faster than iterations on pro, likely due to model size (flash is estimated to be in the 200-400B range).

I found 3.5 pro to be much better than 3.5 flash, but 3.7 flash with high reasoning is comparable and way way faster.

There are no public release of 3.5 pro. Either its a typo, or you have some insider information
Post reply on HN