Live data from Hacker News

Gemini 3.8 Flash and 3.8 Flash Cyber

blog.google

131–140 of 699 posts

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#131

Currently top at https://deepswe.datacurve.ai - beating Opus 5! https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium! Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.

sidenote, but wow sonnet 5 is shockingly bad on this benchmark.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#132
post #113

I've been using Gemini 3.7 for my personal trip planning app. Across multiple benchmarks, it ranks higher on everything I tried: - Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the best at taking a cluster of places and working out a visiting order. - Photo ranking (which photo should be the hero). Gemini can tell whether a photo is of the thing or of the vie…

Gemini 3.7 is my workhorse - fast and good enough for most tasks. Occasionally I go to GPT Sol or Claude to improve Gemini's output or for more complex tasks, but more than of my work usage is Gemini 3.7. Quite happy to test 3.8 now.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#133
post #113

I've been using Gemini 3.7 for my personal trip planning app. Across multiple benchmarks, it ranks higher on everything I tried: - Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the best at taking a cluster of places and working out a visiting order. - Photo ranking (which photo should be the hero). Gemini can tell whether a photo is of the thing or of the vie…

One thing in your comment surprised me: "when a thing opens and closes"

Why you would rely on the model's weights to know opening hours, instead of having the model call a web search tool to verify it on the official site?

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#135

Earlier quoted context omitted.

Google - we're so back

Only 1 point behind the Chinese SOTA from two months ago.

I had qwen 3.8 3bit model drop into chinese on long runs. I had to remind it to use english. Its still better than every gemma model I tried. Gemma deleted files on a harddrive to make space when there was over 2TB free. For long runs, gemma is useless.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#136
post #112

It is more expensive per task than 5.6-sol high: https://artificialanalysis.ai/models/gemini-3-8-flash#price-...

Huh, according to some of those charts, it's both dumber, and more expensive to run against their benchmarking tasks than Fable??? Seems crazy to me.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#137
post #96
post #92

Earlier quoted context omitted.

A fifth of the cost of Opus 5! Google is certainly pushing the completion with this.

Gemini hasn't failed me for personal usage yet. I haven't had the opportunity to use it at work.

I've been using 3.7 Flash to audit the work of Opus High, and Flash finds lots of subtle and insidious defects even while all the unit tests are green.

Then I tell Opus to read the audit report and implement what it agrees with.

Flash is really good at this, and it is blazing fast in Antigravity CLI. Easily 10x faster than Opus.

Can't wait to try 3.8 Flash. If it's good enough, maybe I'll switch Flash to primary and make Opus the auditor.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#138
post #81

Earlier quoted context omitted.

Congratulations, you're this thread's "pelicans are tiresome" comment - it's part of the Hacker News tradition at this point. (Next up is the comment saying that the labs are clearly training for the benchmark.)

The labs are clearly training for the benchmark.

This has been addressed endlessly, for a few years now, and is just as much of a trope as "this benchmark is useless".

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#140
Seems to do reasonably well in opencode according to artificial analysis. https://artificialanalysis.ai/agents/coding-agents

If I had to pay per token I would probably consider using this (they seem to be on the pareto of performance) but not being able to use opencode with a subscription is not really something I'm realistically going to do when claude and codex are around. Also never gotten along well with gemini-cli / antigravity-cli.

Post reply on HN