Live data from Hacker News

Gemini 3.8 Flash and 3.8 Flash Cyber

blog.google

621–630 of 699 posts

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#621
post #418
post #237

Earlier quoted context omitted.

Gemini Flash is the one I get most excited about, because it's so fast and so good at real-world knowledge, and it's improving so fast - look at how much the benchmarks improved in ~1 month. It's just categorically different than anything else. Also, I use it every day, and it just got ~10% better at coding, according to the benchmarks. How is that not exciting?

I use Gemini every day and I've noticed any subjective improvement. In many cases it feels worse because it does fewer Google searches than before. As a result I find it hard to get excited about it. I do think Gemini is underrated on HN though!

[deleted]

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#622
post #462

Earlier quoted context omitted.

the 3.5 pro pretrain was a complete disaster, they shelved it and are now working on gemini 4. 3.0 flash -> 3.8 flash is all post training which is pretty impressive.

Do labs come back from disasters like GDM’s 3.5 pretrain? I am thinking of Meta’s Llama 4. Meta is just now starting to be taken seriously again but they are definitely not at the frontier. And when I say “come back” I mean have an Opus 4.5 moment, which was really mind blowing for me at the time. Fable was a similar leap, just not as big.

Unless the company is going under, why not? Let's say Google releases Gemini Pro 4 tomorrow, and it's better than Fable and Sol; lots of people would switch over to it.

AI models are almost completely interchangeable, so the best/cheapest/fastest whatever will always have a market.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#623
post #113

I've been using Gemini 3.7 for my personal trip planning app. Across multiple benchmarks, it ranks higher on everything I tried: - Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the best at taking a cluster of places and working out a visiting order. - Photo ranking (which photo should be the hero). Gemini can tell whether a photo is of the thing or of the vie…

Gemini 3.7 is my workhorse - fast and good enough for most tasks. Occasionally I go to GPT Sol or Claude to improve Gemini's output or for more complex tasks, but more than of my work usage is Gemini 3.7. Quite happy to test 3.8 now.

What harness do you use for Gemini? Antigravity?

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#624
post #113

I've been using Gemini 3.7 for my personal trip planning app. Across multiple benchmarks, it ranks higher on everything I tried: - Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the best at taking a cluster of places and working out a visiting order. - Photo ranking (which photo should be the hero). Gemini can tell whether a photo is of the thing or of the vie…

> Real world knowledge For awhile now I've found Gemini will use Google search for pretty much any real world knowledge, which is a huge plus IMO. It's basically Google with a much better frontend and no ads/seo nonsense.

There are some subtleties though. For example, with Gemini search grounding (this was a few weeks ago), you cannot give it a domain whitelist, only an excludelist -- with Anthropic's API you can do both. For somewhat niche search tasks like "Find LinkedIn profiles matching this ICP", Anthropic wins there.

That being said, Anthropic is also so insanely expensive for everything I ended up switching that particular part to Exo instead...

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#625
post #113

I've been using Gemini 3.7 for my personal trip planning app. Across multiple benchmarks, it ranks higher on everything I tried: - Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the best at taking a cluster of places and working out a visiting order. - Photo ranking (which photo should be the hero). Gemini can tell whether a photo is of the thing or of the vie…

I've just returned home after doing a 100 day "half lap" of Australia in a camper trailer with my family. I made heavy use of gemini to plan much of it. Getting packed up with a general destination in mind and telling gemini "we're leaving X town at 10am and heading to Y, where should we stop for lunch. My wife has coeliac disease, find us somewhere that does good gluten free options" was one of the many things I regularly leant on it for.

Never had a bad suggestion.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#626

Earlier quoted context omitted.

I setup Luna as main Claude Code driver (so zero anthropic api use) and it nailed crisply a handful of python tasks, gonna continue this way.

Why not use codex or an open source harness?

That is a very valid question. I happen to want to get Claude Code muscle memory under my belt for professional reasons in addition to get side projects advanced, could have settled for Codex else. Also OpenCode with eastern models gets part of the job done. In CC beyond using Luna for the cheap, I am using DeepSeek flash v4 for subagents, that is a further cost shaver. Not sure if in Codex I could do that.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#627
post #200

The speed combined with the fact that this thing is really good at HTML JavaScript is pretty exciting. Here's what I got for 1.8 cents and 13 seconds from the prompt "make me a cool thing in html": https://gisthost.github.io/?6a77bc41a81718c6aaa10d4ab243c59f Transcript here (it was part of a chat): https://gist.github.com/simonw/b6149a49d327164d67d62c3d12992...

it's such a weird split how most AI companies are trying to be the best, but Google really has a different mission statement. they already have users. lots of users. they need to be working on building models they can deploy and use with the most number of people, as they already have the users. i don't know if Gemini models per se are fully is in line with that purpose, but the results we see keep seeming to be in-l…

Google is a top-notch researcher among all the problems people have with it. It has a mix of great products, terrible automated systems (though, I suspect, not as bad as Meta's?) and some historical disappointments

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#628
post #64

Pelicans (thinking effort high, medium, low): https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - high cost 8.9742 cents Here are the 3.7 pelicans for comparison: https://tools.simonwillison.net/markdown-svg-renderer.html?u... - high cost 8.4387 cents (I think thinking level low is a regression on 3.8 compared to 3.7.)

[deleted]

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#629

Earlier quoted context omitted.

>There are numerous benchmarks that measure cost per task, which factors out tokens entirely. Gemini 3.8 flash is significantly lower than Sol on basically all of them https://artificialanalysis.ai/#cost-tabs Not sure if you read your own link but Sol 56 high ranks smack between Gemini 3.8 flash medium and high. Gemini 3.8 flash comes in as more expensive per task than Sol 56 high according to artificial analysis. Lu…

I open the link and I see Flash 3.8 high at 0.58 and Sol at 0.95. I don't understand why you say that "Sol 56 high ranks smack between Gemini 3.8 flash medium and high" but that is clearly wrong.

On cost per intelligence task, Gemini38flash and Sol56 trade back and forth on cost depending on effort level. https://i.imgur.com/zPaWPXx.png As seen in this image, literally: Sol56 high ranks in between Gemini 38 medium and high. The image proves it.

I also included Sol56 xhigh, which ranks above even Gemini38 high.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#630
post #507

Earlier quoted context omitted.

That's the thing. I am completely lost because there are so many redundant paths to get the same thing and I'm trying to figure out which one is the best deal

Well are you looking for a subscription or pay-as-you-go API usage? Subscription? -> Google One plan (http://one.google.com/) API? -> AI Studio (https://aistudio.google.com/) It's not really any different than the choice you'd make with OpenAI/Anthropic depending on how you plan to use it. Except as a hyperscalar, it's also offered first party from Google Cloud (like Claude via Amazon Bedrock or GPT via Microsoft Azu…

good summary, thanks. I used to use Google Cloud for consulting and projects (I worked at Google for a while, and there is some nostalgia) so I have Gemini API via Google Cloud, but I am retired now. Your post reminded me that I need to shut that all down and switch to getting an API key using AI Studio.

I have spent two months experimenting with a wide range of US and Chinese models, and I had a lot of fun doing that, but I am in the process of switching to just using local models, using Gemini on an API if I need it, and once or twice a month when I really need help on something difficult, I use something top-tier like Kimi K3.

Post reply on HN