Live data from Hacker News

Gemini 3.8 Flash and 3.8 Flash Cyber

blog.google

81–90 of 697 posts

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#81
post #64

Pelicans (thinking effort high, medium, low): https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - high cost 8.9742 cents Here are the 3.7 pelicans for comparison: https://tools.simonwillison.net/markdown-svg-renderer.html?u... - high cost 8.4387 cents (I think thinking level low is a regression on 3.8 compared to 3.7.)

I mean no offense but these pelicans are a bit tiresome and a very meaningless benchmark. There's no real difference between any of these svgs across models and model versions anymore.

Congratulations, you're this thread's "pelicans are tiresome" comment - it's part of the Hacker News tradition at this point.

(Next up is the comment saying that the labs are clearly training for the benchmark.)

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#82

Wow this comes after what - 3 or 4 weeks since 3.7 Flash, which was also 3 or 4 weeks after 3.6 Flash IIRC? I eagerly wait more info but sounds like Deepmind without Demis calling the shots has been unleashed and are operating at full speed? Shocker! At this point it is a meme of course, but where is 3.5 Pro :)

A month is not enough time for any meaningful change in an organization the size of Deepmind/Google. These models were surely the result of work streams and teams that started under Demis. I think Demis can safely feel proud Deepmind is getting back on track.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#83

Earlier quoted context omitted.

If by reckless you mean commit, push, deploy without me asking it to, the I agree!

Respectfully: If it's able to deploy without you asking it to, that's a you problem. There are no safeguards?

You need more safeguards for sure, but also it tends to fly off down rabbit holes, rebuilding things in dumb ways, hacking around things, making assumptions etc, it seems very eager to go 'ta da! I did it look how quick I was', sometimes it nails it other times it created a lot of tech debt.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#84
post #66
post #25

Earlier quoted context omitted.

IDK if it's smaller, but I know it's way faster. In one test I did, Flash 3.7 high was ~9.4x faster than Luna High. But, also... Sol crushes Flash 3.7 at writing code in a codebase of any size beyond "tiny". Flash is my go-to for prototyping, and basically anything that isn't writing production code.

Luna is way slow. I don't remember an OpenAI model ever being this slow. edit: I have a subscription; direct call.

Are you using direct or via OpenRouter? I think OpenRouter Luna always uses the `flex` tier, which is quite a bit slower.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#85
post #64

Pelicans (thinking effort high, medium, low): https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - high cost 8.9742 cents Here are the 3.7 pelicans for comparison: https://tools.simonwillison.net/markdown-svg-renderer.html?u... - high cost 8.4387 cents (I think thinking level low is a regression on 3.8 compared to 3.7.)

This is in comparison to Fable: > https://tools.simonwillison.net/markdown-svg-renderer?url=ht... > Took just under 14 minutes to generate, and at 65927 output tokens cost me a hefty $3.30! So 50x cheaper - and how much faster?

The Gemini models have openly trained for SVG output, apparently with a specialism on animals in forms of transport! https://twitter.com/JeffDean/status/2024525132266688757

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#86
post #64

Pelicans (thinking effort high, medium, low): https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - high cost 8.9742 cents Here are the 3.7 pelicans for comparison: https://tools.simonwillison.net/markdown-svg-renderer.html?u... - high cost 8.4387 cents (I think thinking level low is a regression on 3.8 compared to 3.7.)

I mean no offense but these pelicans are a bit tiresome and a very meaningless benchmark. There's no real difference between any of these svgs across models and model versions anymore.

If everyone agreed with you, the comment would disappear near the bottom of the thread

I like the benchmark. Yes, it's near saturation for SotA models, but still quite good to show where smaller models stand in relation to SotA

In this instance, I see a great image, but consistently clipping mudguards (both in 3.8 flash and 3.7 flash)

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#87
post #18

Wait, I didn't realize 3.7 Flash was already beating Sol on a bunch of the benchmarks. Isn't it a way smaller models?

It's pretty good if you can actively steer it , its actually really really good , the antigravity free tier and pro tiers are generous as well . I'm shocked at how fast it generates tokens.

Some would say it's Google's TPUs.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#88
post #81

Earlier quoted context omitted.

I mean no offense but these pelicans are a bit tiresome and a very meaningless benchmark. There's no real difference between any of these svgs across models and model versions anymore.

Congratulations, you're this thread's "pelicans are tiresome" comment - it's part of the Hacker News tradition at this point. (Next up is the comment saying that the labs are clearly training for the benchmark.)

The labs are clearly training for the benchmark.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#89

Earlier quoted context omitted.

If by reckless you mean commit, push, deploy without me asking it to, the I agree!

Respectfully: If it's able to deploy without you asking it to, that's a you problem. There are no safeguards?

Also if it ever says, "I've found the root cause of ..", it definitely has not found the root cause and is making a non evidence based guess as it has run out of ideas.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#90

Currently top at https://deepswe.datacurve.ai - beating Opus 5! https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium! Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.

As of writing this comment, Claude Opus 5 has an intelligence score of 63, not 59 (it's not the same as Gemini 3.8 Flash).

With a score of 59, Gemini 3.8 Flash is in eighth place, falling behind even Grok 4.6, Kimi k3, and GLM 5.3.

https://imgur.com/a/BMOJBED

Post reply on HN