Live data from Hacker News

Gemini 3.8 Flash and 3.8 Flash Cyber

blog.google

71–80 of 697 posts

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#71
post #64

Pelicans (thinking effort high, medium, low): https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - high cost 8.9742 cents Here are the 3.7 pelicans for comparison: https://tools.simonwillison.net/markdown-svg-renderer.html?u... - high cost 8.4387 cents (I think thinking level low is a regression on 3.8 compared to 3.7.)

I mean no offense but these pelicans are a bit tiresome and a very meaningless benchmark. There's no real difference between any of these svgs across models and model versions anymore.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#72
post #64

Pelicans (thinking effort high, medium, low): https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - high cost 8.9742 cents Here are the 3.7 pelicans for comparison: https://tools.simonwillison.net/markdown-svg-renderer.html?u... - high cost 8.4387 cents (I think thinking level low is a regression on 3.8 compared to 3.7.)

This is in comparison to Fable:

> https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

> Took just under 14 minutes to generate, and at 65927 output tokens cost me a hefty $3.30!

So 50x cheaper - and how much faster?

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#73

The flash models, for coding are reckless in my experience. I have a Ultimate subscription, get good quota, but still use Opus 4.6 as it's much more reliable if you manage the context window carefully.

If by reckless you mean commit, push, deploy without me asking it to, the I agree!

Respectfully: If it's able to deploy without you asking it to, that's a you problem. There are no safeguards?

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#74

The flash models, for coding are reckless in my experience. I have a Ultimate subscription, get good quota, but still use Opus 4.6 as it's much more reliable if you manage the context window carefully.

If by reckless you mean commit, push, deploy without me asking it to, the I agree!

It even took my girlfriend on a date, now it prepares for IPO, how do I turn it off?

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#75
post #64

Pelicans (thinking effort high, medium, low): https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - high cost 8.9742 cents Here are the 3.7 pelicans for comparison: https://tools.simonwillison.net/markdown-svg-renderer.html?u... - high cost 8.4387 cents (I think thinking level low is a regression on 3.8 compared to 3.7.)

[deleted]

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#76

Is the Gemini CLI still terrible compared to Claude Code and Codex? The harness the main thing holding back Google models as they could've been the best given all the advantages in compute capacity and training data they initially had, where now even the Google CEO said they're falling behind in agentic tasks, which is sort of a vicious cycle because RLHF relies on human usage.

There is no Gemini CLI anymore, nor you can use Gemini with your own harness unless you pay per-token.

it's called `agy` now

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#77

One place where I find the Flash models surprisingly bad is Google Search's "AI Mode". A recent example - I searched for how to unsubscribe from Pearson emails. Google Search "AI Mode" confidently gave me a sequence of steps along the lines of Settings > Profile > Email preferences > Unsubscribe. Of course, I looked for an unsubscribe link before asking Google. None of those options existed. The correct answer was th…

I think that's just a limitation on the size of the model. I'm pretty sure that they use a pretty small model in those summaries to save money, which naturally makes them a little less smart.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#78

Is the Gemini CLI still terrible compared to Claude Code and Codex? The harness the main thing holding back Google models as they could've been the best given all the advantages in compute capacity and training data they initially had, where now even the Google CEO said they're falling behind in agentic tasks, which is sort of a vicious cycle because RLHF relies on human usage.

In May they replaced the Gemini CLI with the Antigravity CLI.

https://developers.googleblog.com/an-important-update-transi...

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#79

One place where I find the Flash models surprisingly bad is Google Search's "AI Mode". A recent example - I searched for how to unsubscribe from Pearson emails. Google Search "AI Mode" confidently gave me a sequence of steps along the lines of Settings > Profile > Email preferences > Unsubscribe. Of course, I looked for an unsubscribe link before asking Google. None of those options existed. The correct answer was th…

https://www.pearson.com/privacy-center/privacy-notices/full-...

>We will not send marketing emails to a user who has opted out of receiving them. Any marketing communications we send will include an unsubscribe link at the end of the email.

I don't think this is AI's fault. This is Pearson's publishing incorrect information and the only way to really know they are a bunch of lying assholes is to have an account and try to unsubscribe from it.

AI didn't make it up, Pearson's did.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#80

The flash models, for coding are reckless in my experience. I have a Ultimate subscription, get good quota, but still use Opus 4.6 as it's much more reliable if you manage the context window carefully.

I use Flash model as code implementation executor, then have GPT-5.6-Sol or Opus to review the work. Pretty good so far and presumably less expensive.
Post reply on HN