Live data from Hacker News

Gemini 3.8 Flash and 3.8 Flash Cyber

blog.google

461–470 of 699 posts

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#461

Earlier quoted context omitted.

Can you use Pi with a Google Pro AI sub or do you need to use API billing?

Using API billing would be a bummer because, as far as I know, there's no way to set up a spend cap or pre-pay the API key, correct?

They finally allowed for a hard spend cap that's easy to set a few weeks ago: https://aistudio.google.com/spend

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#462

I don't know if Google is having the worst marketing fumble or the most genius marketing one. Their "flash" models are very comparable to other companies' "pro" or "flagship" models. It seems to be a quite counterintuitive naming convention as it undersells the models. Unless they have an even more powerful Gemini Pro in the oven...?

the 3.5 pro pretrain was a complete disaster, they shelved it and are now working on gemini 4. 3.0 flash -> 3.8 flash is all post training which is pretty impressive.

Do labs come back from disasters like GDM’s 3.5 pretrain? I am thinking of Meta’s Llama 4. Meta is just now starting to be taken seriously again but they are definitely not at the frontier. And when I say “come back” I mean have an Opus 4.5 moment, which was really mind blowing for me at the time. Fable was a similar leap, just not as big.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#463

Earlier quoted context omitted.

Easy, have another agent check it. Yeah, I know, just more slop. But I do think the second agent’s eagerness to please is aligned more in your favor in that instance, so it’s likely to find most issues. The bigger problem I’ve found is that it’ll also find all kinds of very minor edge cases that you have to pick through.

Do we add a third one to check the second one which is checking the first? Asking slightly tongue in cheek but at what point does this stop making sense if we can't trust the output, the people creating the models are already getting surprised in bad ways (if we take their words at face value) with how the models are behaving already etc. We have the folks over here saying "AI is amazing" and the other other folks ov…

I mean sure, you can add a third, and a fourth and a fifth one if ur ok with the added cost, latency and it actually helps. Redundancy is a core concept in software and CS and at the heart of making many systems, complex or otherwise, reliable.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#464

Earlier quoted context omitted.

Why are the SVGs getting more detailed rather than just more correct than previous models?

Because people tend to like fidelity more than correctness.

If correctness would matter anymore, people wouldn't be using LLMs in the first place.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#465
post #64

Pelicans (thinking effort high, medium, low): https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - high cost 8.9742 cents Here are the 3.7 pelicans for comparison: https://tools.simonwillison.net/markdown-svg-renderer.html?u... - high cost 8.4387 cents (I think thinking level low is a regression on 3.8 compared to 3.7.)

The rendering of the gullet is very poor, because its both behind the handlebars but in front of the bike frame (impossible geometry). Surprising because gemini is usually pretty good on geo spatial skills. Edit: scrolled down to medium effort, its better but also has a weird clipping issue with the fish in the beak.

LLMs are not intelligent and don't actually understand the concept of a bicycle. Parrots also don't understand human language but they're really good at pretending otherwise.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#466
post #64

Pelicans (thinking effort high, medium, low): https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - high cost 8.9742 cents Here are the 3.7 pelicans for comparison: https://tools.simonwillison.net/markdown-svg-renderer.html?u... - high cost 8.4387 cents (I think thinking level low is a regression on 3.8 compared to 3.7.)

These are becoming unreadable as the reasoning chains expand. I think you should consider reformatting them and either putting the image first or else folding the COT output.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#467
post #404

Wow this comes after what - 3 or 4 weeks since 3.7 Flash, which was also 3 or 4 weeks after 3.6 Flash IIRC? I eagerly wait more info but sounds like Deepmind without Demis calling the shots has been unleashed and are operating at full speed? Shocker! At this point it is a meme of course, but where is 3.5 Pro :)

Now i am awaiting Gemini 3.11 "For Workgroups" to be released early December...

With Gemini 95 following soon after.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#468

Earlier quoted context omitted.

Antigravity is probably the best of the bunch I've tried. I'd say it's pretty comparable to Claude Code (I use both daily).

Antigravity which lacks an auto approve mode? Not really comparable to Claude Code when you're looking to run a team of agents from my experience.

I've curated my auto-approve list to specific commands by approving them with "always allow in this project" (I never want it e.g. committing/pushing to GitHub, removing files, etc without being in the loop) but Antigravity does have both a standard "auto-approve" mode _and_ a "Turbo mode" which disables ALL approvals of all kinds.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#469
post #153

"The knowledge cutoff date for Gemini 3.8 Flash is March 2026 – users can expect updated information for some domains while in others they may experience the model’s knowledge is limited to January 2025 (in line with the Gemini 3 Model Family)." Kind of wild that they haven't (successfully) pretrained a base model since Jan-25.

I'm curious if the knowledge cutoff is important, when the interface (Gemini app) can search online for recent information. Is there a big advantage to having everything internal?

[dead]

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#470
A reminder that google is the only major lab without a meaningful opt-out of training on your data. The only way to opt out is to disable message history entirely, which seems like a darkest of dark patterns to get users to leave "opt in" to training on, because next to nobody wants to use it without message history.
Post reply on HN