Live data from Hacker News

Gemini 3.8 Flash and 3.8 Flash Cyber

blog.google

561–570 of 699 posts

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#561
post #432

Earlier quoted context omitted.

He didn’t write it.

> I asked Claude to fix the grammar of my comment I read this to mean he wrote the comment, then asked Claude to fix the grammar (as many ESL speakers do). Sounds to me like he did write it.

Then I think you read it wrong, because you don't make that mistake unless you copy-paste your comment out of a Claude window and into the comment box.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#562

Currently top at https://deepswe.datacurve.ai - beating Opus 5! https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium! Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.

I've been trying this Gemini 3.8 Flash for a day. Looks not much different than Gemini 3.7 Flash in my use case: I have Codex (gpt-5.6 sol) write up a design plan to implement a feature or refactor a portion of a system I am building, and have Claude (Opus-5) and Gemini (3.8 Flash) review and critique the plan, until all problems are addressed by Codex and approved by the reviewers; then have a cheaper model of Codex…

I find this very interesting, I wonder if there is a public benchmark that reflects this “red team coding critique” aspect of the current SOTA model that reflects what you have observed.

It would be really useful to observe this in a benchmark vs. the more common “go implement this, or fix this bug” type benchmarks that seem to be prevalent.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#563

Earlier quoted context omitted.

As of writing this comment, Claude Opus 5 has an intelligence score of 63, not 59 (it's not the same as Gemini 3.8 Flash). With a score of 59, Gemini 3.8 Flash is in eighth place, falling behind even Grok 4.6, Kimi k3, and GLM 5.3. https://imgur.com/a/BMOJBED

They are all much larger and more expensive models. Google does not have a frontier model right now, but for cheap ones, they are better than event the chinese models now.

> "Google does not have a frontier model right now, but for cheap ones, they are better than event the chinese models now."

Just wow. Someone actually said this.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#564
Gemini 3.7 benchmarks against GDPVal-AA-V2 were 1525 in Aug blog post. However same model against same benchmark is 1482 in today's blog post of Gemini 3.8 release..

Do they make it intentionally to look previous model less superior than current models? or these are the real numbers when re-ran the benchmark..?

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#565
It struck me today using Google Antigravity (Claude Code alternative) just how direct and usably terse Gemini is in prose.

I've complained plenty on here about Claude verbosity and TED-talk phrasing, and it seems by contrast Gemini has already arrived at the dream end-state of Claude from a prose standpoint.

Sometimes I ask for feedback, and I get back a list of multiple-choice options as if it's already ready to go. If I indicate I'm thinking about doing something, sometimes it'll just...do it. (Not in an annoying way.)

It seems very geared toward action in a way that's completely refreshing coming from months steeped in Claude essays.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#566

Earlier quoted context omitted.

I've been trying this Gemini 3.8 Flash for a day. Looks not much different than Gemini 3.7 Flash in my use case: I have Codex (gpt-5.6 sol) write up a design plan to implement a feature or refactor a portion of a system I am building, and have Claude (Opus-5) and Gemini (3.8 Flash) review and critique the plan, until all problems are addressed by Codex and approved by the reviewers; then have a cheaper model of Codex…

I find this very interesting, I wonder if there is a public benchmark that reflects this “red team coding critique” aspect of the current SOTA model that reflects what you have observed. It would be really useful to observe this in a benchmark vs. the more common “go implement this, or fix this bug” type benchmarks that seem to be prevalent.

Yeah, my tool to automate these review loops is https://github.com/wwind123/coding-review-agent-loop . It's basically a script calling Claude, Codex and Antigravity CLI's. The benefit of using CLI's is, the tool uses quota in your subscription plan of these AI providers, which is much cheaper than using extra tokens from the same providers to do the same thing.

A couple of months ago (before opus-5 and gpt-5.6 sol), The ratio of problems caught by codex/claude vs gemini was more like 2:1 to 3:1. But now it seems codex and claude have made huge leaps and gemini is more or less staying put.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#567
so far using Gemini from 2.5 Pro to date (3.7 flash) - the way google trains the model it seems - is to identify top 3 to 4 things to fix first. as a result gemini is not that good in being thorough - but it's a needle mover . Opus5 Opus4.8 and Fable always point out things that Gemini missed. but Gemini was a needle mover - identifying the most important things to fix.

I always enjoy interacting with Gemini . When i ask it questions about designing a new model etc - it's always the most helpful and encouraging . I really want to thank the Google team for this and their happy positive models they generate.

I use gemini flash after a round of deliberation between Sol and Kimi these days on the main plan . Kimi 2.7 paired with Gemini 3.5/4.6/3.7 flash has been my implementer - Kimi k3 and Sol 5.6 have been my planners and code reviewers.

I would have ideally like Anthropic and was on their USD 200 plan - but after they didnt sign the letter for Open source models - i dropped my subscription . Wont make a difference to their lives.

But Sol is great . Combined with Kimi K3 for adversarial plan reviews - you get robust plans . And Sol as a reviewer for Gemini/kimi 2.7 - you get great edge case handling and robust code.

I also integrated muse 1.2 - on the contributor tier - and I Just saw facebook release 1.3 muse . this is great news. The training im permitting is my thanks to FB for releasing the open source models of the past ! Thank you !

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#568

Earlier quoted context omitted.

Nitpick, but in my opinion an LLM is an "it", not a "her" or "he". Using male or female pronouns risks anthropomorphizing them which can lead to unhealthy outcomes.

What about languages, such as Russian, where every single noun has a gender assigned (he, she or it) and AI is a he by default (and everything else is already using pronouns in similar way, like a car is a she, a ship is a he).

That's a good point. Still, "it" makes more sense in English.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#569
It's really disappointing to see social media dismissing gemini so easily.

I think the worst thing we can do is have loyalty towards models. I used to be loyal towards Claude, and my viewpoint changed dramatically when I used codex.

I highly recommend that if you are someone who only used one model so far, that you really give another model a shot and see how it goes. It's very eye opening and gives you a more holistic perspective.

Vendor locking is a big problem when it comes to models, and I hope the software world doesn't do this blindly.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#570

It's really disappointing to see social media dismissing gemini so easily. I think the worst thing we can do is have loyalty towards models. I used to be loyal towards Claude, and my viewpoint changed dramatically when I used codex. I highly recommend that if you are someone who only used one model so far, that you really give another model a shot and see how it goes. It's very eye opening and gives you a more holist…

[dead]
Post reply on HN