Live data from Hacker News

Gemini 3.8 Flash and 3.8 Flash Cyber

blog.google

481–490 of 697 posts

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#481

3.7 flash was by far the best model for image recognition tasks according to my benchmarks. 3.8 flash didn't regress any candidates and improved some specificity (positive ID of common name vs species name of exotic fruit, correct identification of cast/replica of artifact and statue) but is still relatively weaker (26/30) on esoteric public figures (Korean beatboxers). I'm going to have to make my benchmark harder.

I’m very curious about your esoteric public figures benchmark, do you ask it in English or Korean to identify the person? Does it change the result? I wonder if having data labeled in only a given language (or web sources in only a given language) change the output.

I haven't tried asking it in Hangul but these particular artists (and the photos I'm using actually) are linked to their romanized english names on e.g. Fandom so it's not unfindable on the internet

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#482

Earlier quoted context omitted.

Gemini 3.7 is my workhorse - fast and good enough for most tasks. Occasionally I go to GPT Sol or Claude to improve Gemini's output or for more complex tasks, but more than of my work usage is Gemini 3.7. Quite happy to test 3.8 now.

How are you able to get lots of usage out of it cost effectively?

Ultra AI is like $99/month, and it is hard to exhaust unless you are running a lot of concurrent requests.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#483

They've interestingly left out any mention of speed. I have been testing 3.7 flash against 3.5 flash and it seems to lose every time in overall latency. Every benchmark I've seen seems to suggest the opposite[1] - that 3.7 flash is significantly (at times 2x) faster than 3.5 flash - but I have never been able to prove this out in real world use cases. Has anyone found their latency numbers to actually be accurate? Is…

In my niche Redactle puzzle solving benchmark [1] I noticed Gemini 3.8 flash is slightly faster than 3.7 flash. They both smoke every model I've tested. I have not yet run 3.5 flash. Gemini models are great at this task because they seem to have exact Wikipedia text baked into the weights. When I rewrite the wiki text a bit it's not able to one-shot the game so much.

[1]: https://redactle.net/llm-leaderboard

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#484

Earlier quoted context omitted.

Nitpick, but in my opinion an LLM is an "it", not a "her" or "he". Using male or female pronouns risks anthropomorphizing them which can lead to unhealthy outcomes.

FYI in Chinese he/she/it all use the same pronoun "ta" when spoken.

They do have a separate character for animals and objects 它 vs 他/她, though I can imagine learning english and just learning ta = he.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#485
post #439

Earlier quoted context omitted.

Flash models are on the order of 1/10th the size of Opus models, so some flex in the thinking level is fair.

Flash is just a name with no defined or consistent meaning even within labs, let alone between them. Considering both are closed weight, there is no way to truly assess how big the size delta between the two is. Then again, who cares about size, performance and end-to-end speed+cost are what matters along with task adherence, task assessment and so on. Model size also can not be inferred by tokens/sec for a multitude…

It's not totally a mystery

https://arxiv.org/html/2604.24827v1

The short of it is by using hard facts knowledge that is difficult to compress, and then quizzing models on these facts and calibrating against a bunch of open models, you can kind of feel out the size of closed models.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#486

I don't use Gemini, but I thought `cool, let's give this new model a try`. Opened gemini.google.com, and I'm not even surprised. The drop down gives me the following options: - Flash-Lite - 3.6 Flash [new] - 3.1 Pro The above is why i don't use LLM products from Google. If the model is not available right this minute (heck, hours before the release!), then I'm not gonna bother getting back to it tomorrow, because tom…

i have it on gemini app the moment they announced it. pro, student

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#487
post #122
post #115

Earlier quoted context omitted.

"Claude 3.7"?

I asked Claude to fix the grammar of my comment, and it changed "I am using 3.7 for" to "I've been using Claude 3.7", so they sneaked their own name on it.

A two paragraph hacker comment? You burned carbon for that?

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#488

Earlier quoted context omitted.

BTW you're comparing 3.8 flash high to opus 5 medium. 3.8 flash medium scores lower.

Flash models are on the order of 1/10th the size of Opus models, so some flex in the thinking level is fair.

When comparing closed models, the only thing that actually matters to anyone using them is some mix of cost and speed. Considering how much memory a server is using, when evaluating models that you'll never have access to in order to host yourself, doesn't really make sense.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#489
post #314

Earlier quoted context omitted.

As someone who has stubbornly stuck with Claude Code, what's a good harness for Gemini models?

Antigravity is probably the best of the bunch I've tried. I'd say it's pretty comparable to Claude Code (I use both daily).

Just like Claude Code and others it has the same —-dangerously-skip-permissions flag, auto approves everything
Post reply on HN