Live data from Hacker News

Gemini 3.8 Flash and 3.8 Flash Cyber

blog.google

191–200 of 699 posts

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#191
post #96

Earlier quoted context omitted.

Gemini hasn't failed me for personal usage yet. I haven't had the opportunity to use it at work.

I've been using 3.7 Flash to audit the work of Opus High, and Flash finds lots of subtle and insidious defects even while all the unit tests are green. Then I tell Opus to read the audit report and implement what it agrees with. Flash is really good at this, and it is blazing fast in Antigravity CLI. Easily 10x faster than Opus. Can't wait to try 3.8 Flash. If it's good enough, maybe I'll switch Flash to primary and…

Yeah the speed in agy cli is amazing. Whole files get written and "py_compile"d in a single blink of the eye its crazy.

In india, my telco gives me google ai pro for free. And agy with flash goes a long way.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#192

Everyone is censoring models now with anything remotely resembling cyber or bio. I already have problems with my research in mathematical epidemiology because of that - both Sol and Fable simply refuse. They keep pushing people towards Chinese models that can be decensored.

Supposedly Fable 5.1 is better, but I haven't tried it yet. I've run into the same thing with mundane work that is barely bio/cyber adjacent.

Re: Chinese models, even if the model itself isn't censored, some of the big model providers have guardrails now that you can't exceed, which somewhat defeats the purpose.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#193
post #41

Currently top at https://deepswe.datacurve.ai - beating Opus 5! https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium! Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.

Crushing it on DeepSWE is a very big deal. Excited to give this a try.

I know everyone is benchmaxxing but this one feels one step too far. Doesn't DeepSWE have both public and private tasks? I'd love to see the diff here.

It looks more like Google execs losing their mind and pressuring researchers to put DeepSWE directly into the training set.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#194

I don't use Gemini, but I thought `cool, let's give this new model a try`. Opened gemini.google.com, and I'm not even surprised. The drop down gives me the following options: - Flash-Lite - 3.6 Flash [new] - 3.1 Pro The above is why i don't use LLM products from Google. If the model is not available right this minute (heck, hours before the release!), then I'm not gonna bother getting back to it tomorrow, because tom…

It's such a weird attitude, especially considering that 1) it's readily available on AI Studio 2) Anthropic models were not always available the moment they got released either. (It also shows that the internet isn't dead. Even people who are not aware of Google AI Studio can express their valuable opinions on LLMs!)

"The new Gemini model isn't available in Gemini, the Gemini App Gemini model is two versions behind and marked as new and the actual new model is in AI Studio" is the kind of problem only Google has though.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#195

I don't use Gemini, but I thought `cool, let's give this new model a try`. Opened gemini.google.com, and I'm not even surprised. The drop down gives me the following options: - Flash-Lite - 3.6 Flash [new] - 3.1 Pro The above is why i don't use LLM products from Google. If the model is not available right this minute (heck, hours before the release!), then I'm not gonna bother getting back to it tomorrow, because tom…

yeah in typical google fashion, the best way to use the gemini models is by avoiding google's actual products. i've got a vision project where gemini flash is the best option by a long shot, and i just use openrouter so i don't have to navigate google's mess.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#196
post #96

Earlier quoted context omitted.

Gemini hasn't failed me for personal usage yet. I haven't had the opportunity to use it at work.

I've been using 3.7 Flash to audit the work of Opus High, and Flash finds lots of subtle and insidious defects even while all the unit tests are green. Then I tell Opus to read the audit report and implement what it agrees with. Flash is really good at this, and it is blazing fast in Antigravity CLI. Easily 10x faster than Opus. Can't wait to try 3.8 Flash. If it's good enough, maybe I'll switch Flash to primary and…

it's very fast but it still doesnt come close to 5.6 sol, at least for me, in terms of gathering the context necessary to do extensive changes.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#197
post #64

Pelicans (thinking effort high, medium, low): https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - high cost 8.9742 cents Here are the 3.7 pelicans for comparison: https://tools.simonwillison.net/markdown-svg-renderer.html?u... - high cost 8.4387 cents (I think thinking level low is a regression on 3.8 compared to 3.7.)

The rendering of the gullet is very poor, because its both behind the handlebars but in front of the bike frame (impossible geometry). Surprising because gemini is usually pretty good on geo spatial skills.

Edit: scrolled down to medium effort, its better but also has a weird clipping issue with the fish in the beak.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#198
post #153

"The knowledge cutoff date for Gemini 3.8 Flash is March 2026 – users can expect updated information for some domains while in others they may experience the model’s knowledge is limited to January 2025 (in line with the Gemini 3 Model Family)." Kind of wild that they haven't (successfully) pretrained a base model since Jan-25.

I'm curious if the knowledge cutoff is important, when the interface (Gemini app) can search online for recent information. Is there a big advantage to having everything internal?

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#199
post #94
post #56

Earlier quoted context omitted.

Not just speed, also reliability. IME, Gemini's speed and quality doesn't degrade badly during weekday working hours compared to OAI, and especially Anthropic.

I've had Gemini model API use degrade the most out of OAI/Anthropic/Google (often "over capacity" vs true failures) Not sure on consumer/product use though

That's interesting to hear. I should have added that I use Gemini through Google AI Studio as my general chat model, which probably explains our wildly different experiences.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#200
The speed combined with the fact that this thing is really good at HTML JavaScript is pretty exciting.

Here's what I got for 1.8 cents and 13 seconds from the prompt "make me a cool thing in html":

https://gisthost.github.io/?6a77bc41a81718c6aaa10d4ab243c59f

Transcript here (it was part of a chat): https://gist.github.com/simonw/b6149a49d327164d67d62c3d12992...

Post reply on HN