Live data from Hacker News

Gemini 3.8 Flash and 3.8 Flash Cyber

blog.google

61–70 of 697 posts

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#61

The flash models, for coding are reckless in my experience. I have a Ultimate subscription, get good quota, but still use Opus 4.6 as it's much more reliable if you manage the context window carefully.

> The flash models, for coding are reckless in my experience.

My experience is that antigravity is awful and reckless - but that the model itself isn't.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#62

The flash models, for coding are reckless in my experience. I have a Ultimate subscription, get good quota, but still use Opus 4.6 as it's much more reliable if you manage the context window carefully.

If by reckless you mean commit, push, deploy without me asking it to, the I agree!

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#63

Is the Gemini CLI still terrible compared to Claude Code and Codex? The harness the main thing holding back Google models as they could've been the best given all the advantages in compute capacity and training data they initially had, where now even the Google CEO said they're falling behind in agentic tasks, which is sort of a vicious cycle because RLHF relies on human usage.

It was superseded by the antigravity CLI.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#64
Pelicans (thinking effort high, medium, low): https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - high cost 8.9742 cents

Here are the 3.7 pelicans for comparison: https://tools.simonwillison.net/markdown-svg-renderer.html?u... - high cost 8.4387 cents

(I think thinking level low is a regression on 3.8 compared to 3.7.)

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#65

Is the Gemini CLI still terrible compared to Claude Code and Codex? The harness the main thing holding back Google models as they could've been the best given all the advantages in compute capacity and training data they initially had, where now even the Google CEO said they're falling behind in agentic tasks, which is sort of a vicious cycle because RLHF relies on human usage.

That was sunset and replaced by Antigravity. FWIW until I abandoned it knowing the sunsetting, I was able to get good behavior out of Gemini CLI with overriding the system prompt. The default prompt crippled the harness with very poor instructions, but there was a hidden ENV to override it. Replacing it with Claude Code like prompts based on the model selected, it ran at a much higher intelligence level full stack with significantly less errors.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#66
post #25
post #18

Wait, I didn't realize 3.7 Flash was already beating Sol on a bunch of the benchmarks. Isn't it a way smaller models?

IDK if it's smaller, but I know it's way faster. In one test I did, Flash 3.7 high was ~9.4x faster than Luna High. But, also... Sol crushes Flash 3.7 at writing code in a codebase of any size beyond "tiny". Flash is my go-to for prototyping, and basically anything that isn't writing production code.

Luna is way slow. I don't remember an OpenAI model ever being this slow.

edit: I have a subscription; direct call.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#68
post #34

Is there any reason to even use 3.1 Pro now?

It is still going to be better at text work, skills, document review, deep reasoning, architecture review, etc. It is only 6 months old, it isn’t like its world knowledge and software knowledge is really out of date. Use it to churn on harder design problems.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#69

Currently top at https://deepswe.datacurve.ai - beating Opus 5! https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium! Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.

Google - we're so back

Only 1 point behind the Chinese SOTA from two months ago.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#70
One place where I find the Flash models surprisingly bad is Google Search's "AI Mode".

A recent example - I searched for how to unsubscribe from Pearson emails. Google Search "AI Mode" confidently gave me a sequence of steps along the lines of Settings > Profile > Email preferences > Unsubscribe.

Of course, I looked for an unsubscribe link before asking Google. None of those options existed. The correct answer was there is no way to unsubscribe through the account, so I just blockthe emails instead.

I've run into this pattern quite a few times. AI Mode seems to make up things all the time.

Post reply on HN