3.7 flash was by far the best model for image recognition tasks according to my benchmarks. 3.8 flash didn't regress any candidates and improved some specificity (positive ID of common name vs species name of exotic fruit, correct identification of cast/replica of artifact and statue) but is still relatively weaker (26/30) on esoteric public figures (Korean beatboxers). I'm going to have to make my benchmark harder.
I’m very curious about your esoteric public figures benchmark, do you ask it in English or Korean to identify the person? Does it change the result? I wonder if having data labeled in only a given language (or web sources in only a given language) change the output.
Gemini 3.8 Flash and 3.8 Flash Cyber
481–490 of 699 posts
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#482Earlier quoted context omitted.
Gemini 3.7 is my workhorse - fast and good enough for most tasks. Occasionally I go to GPT Sol or Claude to improve Gemini's output or for more complex tasks, but more than of my work usage is Gemini 3.7. Quite happy to test 3.8 now.
How are you able to get lots of usage out of it cost effectively?
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#483They've interestingly left out any mention of speed. I have been testing 3.7 flash against 3.5 flash and it seems to lose every time in overall latency. Every benchmark I've seen seems to suggest the opposite[1] - that 3.7 flash is significantly (at times 2x) faster than 3.5 flash - but I have never been able to prove this out in real world use cases. Has anyone found their latency numbers to actually be accurate? Is…
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#484Earlier quoted context omitted.
Nitpick, but in my opinion an LLM is an "it", not a "her" or "he". Using male or female pronouns risks anthropomorphizing them which can lead to unhealthy outcomes.
FYI in Chinese he/she/it all use the same pronoun "ta" when spoken.
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#485Earlier quoted context omitted.
Flash models are on the order of 1/10th the size of Opus models, so some flex in the thinking level is fair.
Flash is just a name with no defined or consistent meaning even within labs, let alone between them. Considering both are closed weight, there is no way to truly assess how big the size delta between the two is. Then again, who cares about size, performance and end-to-end speed+cost are what matters along with task adherence, task assessment and so on. Model size also can not be inferred by tokens/sec for a multitude…
https://arxiv.org/html/2604.24827v1
The short of it is by using hard facts knowledge that is difficult to compress, and then quizzing models on these facts and calibrating against a bunch of open models, you can kind of feel out the size of closed models.
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#486I don't use Gemini, but I thought `cool, let's give this new model a try`. Opened gemini.google.com, and I'm not even surprised. The drop down gives me the following options: - Flash-Lite - 3.6 Flash [new] - 3.1 Pro The above is why i don't use LLM products from Google. If the model is not available right this minute (heck, hours before the release!), then I'm not gonna bother getting back to it tomorrow, because tom…
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#487Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#488Earlier quoted context omitted.
BTW you're comparing 3.8 flash high to opus 5 medium. 3.8 flash medium scores lower.
Flash models are on the order of 1/10th the size of Opus models, so some flex in the thinking level is fair.
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#489Earlier quoted context omitted.
As someone who has stubbornly stuck with Claude Code, what's a good harness for Gemini models?
Antigravity is probably the best of the bunch I've tried. I'd say it's pretty comparable to Claude Code (I use both daily).