Earlier quoted context omitted.
He didn’t write it.
> I asked Claude to fix the grammar of my comment I read this to mean he wrote the comment, then asked Claude to fix the grammar (as many ESL speakers do). Sounds to me like he did write it.
Gemini 3.8 Flash and 3.8 Flash Cyber
561–570 of 699 posts
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#562Currently top at https://deepswe.datacurve.ai - beating Opus 5! https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium! Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.
I've been trying this Gemini 3.8 Flash for a day. Looks not much different than Gemini 3.7 Flash in my use case: I have Codex (gpt-5.6 sol) write up a design plan to implement a feature or refactor a portion of a system I am building, and have Claude (Opus-5) and Gemini (3.8 Flash) review and critique the plan, until all problems are addressed by Codex and approved by the reviewers; then have a cheaper model of Codex…
It would be really useful to observe this in a benchmark vs. the more common “go implement this, or fix this bug” type benchmarks that seem to be prevalent.
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#563Earlier quoted context omitted.
As of writing this comment, Claude Opus 5 has an intelligence score of 63, not 59 (it's not the same as Gemini 3.8 Flash). With a score of 59, Gemini 3.8 Flash is in eighth place, falling behind even Grok 4.6, Kimi k3, and GLM 5.3. https://imgur.com/a/BMOJBED
They are all much larger and more expensive models. Google does not have a frontier model right now, but for cheap ones, they are better than event the chinese models now.
Just wow. Someone actually said this.
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#564Do they make it intentionally to look previous model less superior than current models? or these are the real numbers when re-ran the benchmark..?
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#565I've complained plenty on here about Claude verbosity and TED-talk phrasing, and it seems by contrast Gemini has already arrived at the dream end-state of Claude from a prose standpoint.
Sometimes I ask for feedback, and I get back a list of multiple-choice options as if it's already ready to go. If I indicate I'm thinking about doing something, sometimes it'll just...do it. (Not in an annoying way.)
It seems very geared toward action in a way that's completely refreshing coming from months steeped in Claude essays.
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#566Earlier quoted context omitted.
I've been trying this Gemini 3.8 Flash for a day. Looks not much different than Gemini 3.7 Flash in my use case: I have Codex (gpt-5.6 sol) write up a design plan to implement a feature or refactor a portion of a system I am building, and have Claude (Opus-5) and Gemini (3.8 Flash) review and critique the plan, until all problems are addressed by Codex and approved by the reviewers; then have a cheaper model of Codex…
I find this very interesting, I wonder if there is a public benchmark that reflects this “red team coding critique” aspect of the current SOTA model that reflects what you have observed. It would be really useful to observe this in a benchmark vs. the more common “go implement this, or fix this bug” type benchmarks that seem to be prevalent.
A couple of months ago (before opus-5 and gpt-5.6 sol), The ratio of problems caught by codex/claude vs gemini was more like 2:1 to 3:1. But now it seems codex and claude have made huge leaps and gemini is more or less staying put.
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#567I always enjoy interacting with Gemini . When i ask it questions about designing a new model etc - it's always the most helpful and encouraging . I really want to thank the Google team for this and their happy positive models they generate.
I use gemini flash after a round of deliberation between Sol and Kimi these days on the main plan . Kimi 2.7 paired with Gemini 3.5/4.6/3.7 flash has been my implementer - Kimi k3 and Sol 5.6 have been my planners and code reviewers.
I would have ideally like Anthropic and was on their USD 200 plan - but after they didnt sign the letter for Open source models - i dropped my subscription . Wont make a difference to their lives.
But Sol is great . Combined with Kimi K3 for adversarial plan reviews - you get robust plans . And Sol as a reviewer for Gemini/kimi 2.7 - you get great edge case handling and robust code.
I also integrated muse 1.2 - on the contributor tier - and I Just saw facebook release 1.3 muse . this is great news. The training im permitting is my thanks to FB for releasing the open source models of the past ! Thank you !
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#568Earlier quoted context omitted.
Nitpick, but in my opinion an LLM is an "it", not a "her" or "he". Using male or female pronouns risks anthropomorphizing them which can lead to unhealthy outcomes.
What about languages, such as Russian, where every single noun has a gender assigned (he, she or it) and AI is a he by default (and everything else is already using pronouns in similar way, like a car is a she, a ship is a he).
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#569I think the worst thing we can do is have loyalty towards models. I used to be loyal towards Claude, and my viewpoint changed dramatically when I used codex.
I highly recommend that if you are someone who only used one model so far, that you really give another model a shot and see how it goes. It's very eye opening and gives you a more holistic perspective.
Vendor locking is a big problem when it comes to models, and I hope the software world doesn't do this blindly.
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#570It's really disappointing to see social media dismissing gemini so easily. I think the worst thing we can do is have loyalty towards models. I used to be loyal towards Claude, and my viewpoint changed dramatically when I used codex. I highly recommend that if you are someone who only used one model so far, that you really give another model a shot and see how it goes. It's very eye opening and gives you a more holist…