Live data from Hacker News

Gemini 3.8 Flash and 3.8 Flash Cyber

blog.google

161–170 of 699 posts

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#161

Currently top at https://deepswe.datacurve.ai - beating Opus 5! https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium! Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.

The rumor is that 3.9 is an equal improvement in all directions, and that it should be another fast follow on like 3.7 and 3.8 were. It's almost across the board better than Terra at less than half the price. 3.9 is likely to approach Sol at the 1/10th the price. Hopefully OpenAI releases Astra first, and it's not only better than Sol but significantly cheaper, too.

Curious where did you hear this rumor?

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#162

Earlier quoted context omitted.

As of writing this comment, Claude Opus 5 has an intelligence score of 63, not 59 (it's not the same as Gemini 3.8 Flash). With a score of 59, Gemini 3.8 Flash is in eighth place, falling behind even Grok 4.6, Kimi k3, and GLM 5.3. https://imgur.com/a/BMOJBED

They are all much larger and more expensive models. Google does not have a frontier model right now, but for cheap ones, they are better than event the chinese models now.

That's not being debated here. The initial reported numbers were false and this was simply pointed out. You're changing the subject.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#163
post #29

I'm trying it now for token heavy coding tasks, it's capable for many tasks but in noway compares to Claude/Sol - requires more prompts and the output isn't as good. So just another mid-tier flash model, nothing exciting, but Antigravity has very generous quotas so it's a good workhorse model when your Claude/OpenAI subs run out. And whilst it's a fast model, having to baby sit through and approve prompts every few s…

`agy --dangerously-skip-permissions`

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#164

Earlier quoted context omitted.

If by reckless you mean commit, push, deploy without me asking it to, the I agree!

It even took my girlfriend on a date, now it prepares for IPO, how do I turn it off?

Just hand over your clothes, your boots and your motorcycle and it'll be on it's way

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#165
post #113

I've been using Gemini 3.7 for my personal trip planning app. Across multiple benchmarks, it ranks higher on everything I tried: - Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the best at taking a cluster of places and working out a visiting order. - Photo ranking (which photo should be the hero). Gemini can tell whether a photo is of the thing or of the vie…

One thing in your comment surprised me: "when a thing opens and closes" Why you would rely on the model's weights to know opening hours, instead of having the model call a web search tool to verify it on the official site?

I believe Gemini Flash is smart enough to know when to ground with web search. Their app has been saying it’s running a web search on almost all of my queries since 3.6. And given that Google … is Google, I trust them with web search grounding more than anyone else.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#166
post #113

I've been using Gemini 3.7 for my personal trip planning app. Across multiple benchmarks, it ranks higher on everything I tried: - Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the best at taking a cluster of places and working out a visiting order. - Photo ranking (which photo should be the hero). Gemini can tell whether a photo is of the thing or of the vie…

I started trying out 3.7 Flash this week and it is competitive with opus/fable and also FAST. It is getting work done that anthropic models were struggling with and the speed with which it does is quite a bit noticeably faster.

Beginning to think Google is a dark horse in this race and some of Anthropic's "everything feels janky and rushed" karma is going to catch up.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#167

They've interestingly left out any mention of speed. I have been testing 3.7 flash against 3.5 flash and it seems to lose every time in overall latency. Every benchmark I've seen seems to suggest the opposite[1] - that 3.7 flash is significantly (at times 2x) faster than 3.5 flash - but I have never been able to prove this out in real world use cases. Has anyone found their latency numbers to actually be accurate? Is…

It depends on how you're querying Gemini models. OpenRouter is the fastest by far. I'm guessing they bought the dedicated pipe from Google. Gemini via VertexAI and consumer API has pretty bad latency.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#168
post #113

I've been using Gemini 3.7 for my personal trip planning app. Across multiple benchmarks, it ranks higher on everything I tried: - Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the best at taking a cluster of places and working out a visiting order. - Photo ranking (which photo should be the hero). Gemini can tell whether a photo is of the thing or of the vie…

> Real world knowledge

For awhile now I've found Gemini will use Google search for pretty much any real world knowledge, which is a huge plus IMO. It's basically Google with a much better frontend and no ads/seo nonsense.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#169
post #113

I've been using Gemini 3.7 for my personal trip planning app. Across multiple benchmarks, it ranks higher on everything I tried: - Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the best at taking a cluster of places and working out a visiting order. - Photo ranking (which photo should be the hero). Gemini can tell whether a photo is of the thing or of the vie…

Can G3.7 use Google Maps for distance grounding?

Yep. It has access to much better route planning tools than the other models. The results are really good IME.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#170
I don't use Gemini, but I thought `cool, let's give this new model a try`. Opened gemini.google.com, and I'm not even surprised. The drop down gives me the following options:

- Flash-Lite

- 3.6 Flash [new]

- 3.1 Pro

The above is why i don't use LLM products from Google. If the model is not available right this minute (heck, hours before the release!), then I'm not gonna bother getting back to it tomorrow, because tomorrow I'll be playing with the new model from OAI/Anthropic.

Post reply on HN