Live data from Hacker News

Gemini 3.8 Flash and 3.8 Flash Cyber

blog.google

231–240 of 697 posts

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#231
post #162

Earlier quoted context omitted.

That's not being debated here. The initial reported numbers were false and this was simply pointed out. You're changing the subject.

Opus 5 medium has the same score as 3.8 flash on artificial analysis intelligence index. Are you implying Google or Artificial Analysis are reporting false numbers? What's your source?

BTW you're comparing 3.8 flash high to opus 5 medium. 3.8 flash medium scores lower.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#232

I don't use Gemini, but I thought `cool, let's give this new model a try`. Opened gemini.google.com, and I'm not even surprised. The drop down gives me the following options: - Flash-Lite - 3.6 Flash [new] - 3.1 Pro The above is why i don't use LLM products from Google. If the model is not available right this minute (heck, hours before the release!), then I'm not gonna bother getting back to it tomorrow, because tom…

It's such a weird attitude, especially considering that 1) it's readily available on AI Studio 2) Anthropic models were not always available the moment they got released either. (It also shows that the internet isn't dead. Even people who are not aware of Google AI Studio can express their valuable opinions on LLMs!)

> it's readily available on AI Studio

AI Studio? Seriously, the hell is that? Gemini, AI Studio, Antigravity - what is all that nonsense? The 3.8 Flash announcement says the model is available to Google AI Pro customers. Is it the same as Gemini Pro, or some sort of AI Studio Pro? Based on the comments, i see the model is available in the Gemini App, not available in the UI, not available to Workspace accounts but is available to some personal accounts, yet I'm not a Workspace user. Some people have already mentioned that they are paid customers, yet they don't see the new model.

I know Google loves asking graph problems during their tech interviews, but I can't wrap my head why the customers should solve these problems as well.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#233

I don't know if Google is having the worst marketing fumble or the most genius marketing one. Their "flash" models are very comparable to other companies' "pro" or "flagship" models. It seems to be a quite counterintuitive naming convention as it undersells the models. Unless they have an even more powerful Gemini Pro in the oven...?

Well if that's the case, it's been in the oven for quite a while now.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#234

Earlier quoted context omitted.

Gemini 3.7 is my workhorse - fast and good enough for most tasks. Occasionally I go to GPT Sol or Claude to improve Gemini's output or for more complex tasks, but more than of my work usage is Gemini 3.7. Quite happy to test 3.8 now.

Same here. I see so many people obsessing over the latest most state of the art bleeding edge models and yelling at Google for not being there, but I feel like the vast majority of people don't actually need those models. Flash has just been super useful and incredibly fast in my experience.

I prefer luna for most development, especially when I am guiding the process. Sometimes terra. I have had terrible results coding with sol. It is way over-tuned on RL to make something that completes the task, no matter what. I end up with way too much code that does a lot of things I didn't ask for.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#235

Not to rain on anyone's parade but I find it strange how excited and giddy people on HN get for any new X.X model releases. Pumping it straight to the top, clamoring to use it, check and compare benchmarks, bragging about it being your "daily driver"? Are you people truly this excited about this crap? I mean I guess if you work for Google or Anthropic or whatever I could see it??? Otherwise, are these just bot commen…

Given that they push capabilities at the pareto frontier, yeah

A lot of us use these in our services, so we're getting an upgrade "for free"

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#236
post #124

The biggest problem with Gemini is that its performance degrades the longer you use it for coding. Is it just me?

Seems maybe you’re keeping a forever-session and multiple independent tasks end up overstaying in context?

I would say either start new sessions for new tasks or limit the context to something smaller than 1M.

I usually start with research/planning session, this goes into a detailed implementation plan and then a new session for the actual implementation.

If it's complex problem maybe a review/adversarial step between plan and implementation.

Also with forever-session any time you take a longer break (depends on model and provider as to how long) you will push an entire big context again without caching even if you don't need it. With 1M context this gets expensive.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#237

Not to rain on anyone's parade but I find it strange how excited and giddy people on HN get for any new X.X model releases. Pumping it straight to the top, clamoring to use it, check and compare benchmarks, bragging about it being your "daily driver"? Are you people truly this excited about this crap? I mean I guess if you work for Google or Anthropic or whatever I could see it??? Otherwise, are these just bot commen…

Gemini Flash is the one I get most excited about, because it's so fast and so good at real-world knowledge, and it's improving so fast - look at how much the benchmarks improved in ~1 month. It's just categorically different than anything else.

Also, I use it every day, and it just got ~10% better at coding, according to the benchmarks. How is that not exciting?

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#238
post #153

"The knowledge cutoff date for Gemini 3.8 Flash is March 2026 – users can expect updated information for some domains while in others they may experience the model’s knowledge is limited to January 2025 (in line with the Gemini 3 Model Family)." Kind of wild that they haven't (successfully) pretrained a base model since Jan-25.

I'm curious if the knowledge cutoff is important, when the interface (Gemini app) can search online for recent information. Is there a big advantage to having everything internal?

Not directly - but latest research advancements, cleaner / richer datasets, etc. still require fresh base models. Not everything can be fixed through post training alone (e.g. why GPT-5.5 "Spud" was such a big jump, and also why GPT-6 "Astra" is now supposedly another big leap). Ofc model size etc also plays a role, but my (admittedly limited) understanding is that new base models _can_ also lead to big jumps even keeping parameter counts constant.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#239
post #64

Pelicans (thinking effort high, medium, low): https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - high cost 8.9742 cents Here are the 3.7 pelicans for comparison: https://tools.simonwillison.net/markdown-svg-renderer.html?u... - high cost 8.4387 cents (I think thinking level low is a regression on 3.8 compared to 3.7.)

Why are the SVGs getting more detailed rather than just more correct than previous models?

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#240
post #29

I'm trying it now for token heavy coding tasks, it's capable for many tasks but in noway compares to Claude/Sol - requires more prompts and the output isn't as good. So just another mid-tier flash model, nothing exciting, but Antigravity has very generous quotas so it's a good workhorse model when your Claude/OpenAI subs run out. And whilst it's a fast model, having to baby sit through and approve prompts every few s…

`agy --dangerously-skip-permissions`

anyway to do this with the Antigravity macOS App?
Post reply on HN