Live data from Hacker News

Gemini 3.8 Flash and 3.8 Flash Cyber

blog.google

111–120 of 699 posts

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#111
post #44

Is the google infra stable enough right now? At the start of the year, the flash model was unusable for a whole month via gemini CLI. They could not fix it for a whole month and I was a paid customer.

I use it quite a lot and after a week of use I’m being hard rate limited

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#113
I've been using Gemini 3.7 for my personal trip planning app. Across multiple benchmarks, it ranks higher on everything I tried:

- Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the best at taking a cluster of places and working out a visiting order.

- Photo ranking (which photo should be the hero). Gemini can tell whether a photo is of the thing or of the view from it.

- Document parsing (extracting the relevant trip info from PDFs).

If you use LLMs for anything other than coding, I definitely recommend not discounting Gemini like I did just because other models are more popular.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#114
post #101

Looks like the strategy of regular updates with incremental improvements is working out well. Interestingly, the biggest jump in Artificial Analysis Intelligence Index score is for reasoning level Medium ( 3.7 was 51, 53, 57 for Low, Medium and High, 3.8 is 52,57, 59 respectively). I think scores at lower reasoning levels are more indicative of model capability since higher reasoning levels are focussed on benchmaxxi…

I'm not an expert but I agree with your statement on the lower reasoning levels.

Lots of models seem to just allow the model to "bloatmax" tokens in order to get bumps at high/max reasoning levels. Many of the max reasoning levels allow models to use up to double or more the tokens the next lowest reasoning level uses. Its basically only useful for people who have no cost or time stipulations on anything.

I think I actually preferred it when we had models that either had reasoning enabled or didn't.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#115
post #113

I've been using Gemini 3.7 for my personal trip planning app. Across multiple benchmarks, it ranks higher on everything I tried: - Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the best at taking a cluster of places and working out a visiting order. - Photo ranking (which photo should be the hero). Gemini can tell whether a photo is of the thing or of the vie…

"Claude 3.7"?

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#120

Earlier quoted context omitted.

On artificial analysis it's only equal to opus 5 medium effort. Opus 5 max scores 63. Further, opus 5 medium outputs 4x fewer tokens to achieve the same result, negating a lot of the speed difference.

A comparison to an artificial score and a comparison to “the same task” These folks must laugh themselves to sleep. This whole industry hoodwinked the masses. It’s impressive.

It’s all just vibes
Post reply on HN