Live data from Hacker News

Gemini 3.8 Flash and 3.8 Flash Cyber

blog.google

571–580 of 699 posts

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#573

Earlier quoted context omitted.

Gemini 3.7 is my workhorse - fast and good enough for most tasks. Occasionally I go to GPT Sol or Claude to improve Gemini's output or for more complex tasks, but more than of my work usage is Gemini 3.7. Quite happy to test 3.8 now.

Same here. I see so many people obsessing over the latest most state of the art bleeding edge models and yelling at Google for not being there, but I feel like the vast majority of people don't actually need those models. Flash has just been super useful and incredibly fast in my experience.

And much better bang for the buck, as well.

When I read all the issues people have with Claude - the aggressive guardrails, the cost, how quickly it burns tokens - it seems almost masochistic to use it. Just seems like herd behavior - people use it because everyone else is using it, and because they believe it’s the “best”, whatever that means. (Benchmarks certainly don’t help define that.)

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#574
post #363

Earlier quoted context omitted.

There are numerous benchmarks that measure cost per task, which factors out tokens entirely. Gemini 3.8 flash is significantly lower than Sol on basically all of them https://artificialanalysis.ai/#cost-tabs That said, Luna is the undisputed king here at the moment and is what I use as my workhorse model.

>There are numerous benchmarks that measure cost per task, which factors out tokens entirely. Gemini 3.8 flash is significantly lower than Sol on basically all of them https://artificialanalysis.ai/#cost-tabs Not sure if you read your own link but Sol 56 high ranks smack between Gemini 3.8 flash medium and high. Gemini 3.8 flash comes in as more expensive per task than Sol 56 high according to artificial analysis. Lu…

I open the link and I see Flash 3.8 high at 0.58 and Sol at 0.95. I don't understand why you say that "Sol 56 high ranks smack between Gemini 3.8 flash medium and high" but that is clearly wrong.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#575

Earlier quoted context omitted.

Antigravity is probably the best of the bunch I've tried. I'd say it's pretty comparable to Claude Code (I use both daily).

Antigravity which lacks an auto approve mode? Not really comparable to Claude Code when you're looking to run a team of agents from my experience.

[flagged]

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#576
post #509
post #113

I've been using Gemini 3.7 for my personal trip planning app. Across multiple benchmarks, it ranks higher on everything I tried: - Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the best at taking a cluster of places and working out a visiting order. - Photo ranking (which photo should be the hero). Gemini can tell whether a photo is of the thing or of the vie…

In my experience, Gemini 3.7 is excellent for general non-coding tasks. But for coding, especially backend development, I still find models like Opus 5 and GPT-5.6 more reliable.

I’ve been using Gemini for coding daily since 3.5, mostly on some pretty complex ML engineering projects. It’s great.

What’s an example of not being “reliable”?

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#577

Currently top at https://deepswe.datacurve.ai - beating Opus 5! https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium! Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.

The rumor is that 3.9 is an equal improvement in all directions, and that it should be another fast follow on like 3.7 and 3.8 were. It's almost across the board better than Terra at less than half the price. 3.9 is likely to approach Sol at the 1/10th the price. Hopefully OpenAI releases Astra first, and it's not only better than Sol but significantly cheaper, too.

Reddit thinks Astra will be released today (Thursday/Friday)

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#578

Earlier quoted context omitted.

Because intent supposes will which supposes consciousness, and these aren't.

I’m convinced consciousness isn’t the special thing we think it is.

If consciousness isn't a special thing, then arguing that LLM parameters are conscious is panpsychism or any control loop architecture that observes the outside world, updates an internal state and produces an observable action is considered conscious.

In both cases, LLMs are just as boring as the consciousness definition.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#579
post #314
post #309

People have been sleeping on Gemini lately but these last few Flash releases (which were very rapid) are damn good. These sort of fast and cheap models are great for tasks that are verifiable and can be retried infinitely (like coding), you can basically get frontier results with a good harness (at a fraction of the time and money).

As someone who has stubbornly stuck with Claude Code, what's a good harness for Gemini models?

If you're on their subscription plan - agy cli or antigravity ui is the only choice i think.

Anyway - if you're a dev - you would be writing your own agentic env right ? that's the best way forward. I wont tell you more than this . but if you're not - you are losing out .

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#580
post #363

Earlier quoted context omitted.

I honestly can't believe serious people are making this argument on a straight face. Gemini 3.7 flash outputs so many tokens per answer it doesn't matter how fast its TPS is, sol will end up being both cheaper and faster than Gemini. So ppl are paying more for a given task, waiting longer and using a dumber intelligence because "TPS number shiny". Gemini 3.8 outputs 11k more tokens PER TASK on average in AAII than 3.…

There are numerous benchmarks that measure cost per task, which factors out tokens entirely. Gemini 3.8 flash is significantly lower than Sol on basically all of them https://artificialanalysis.ai/#cost-tabs That said, Luna is the undisputed king here at the moment and is what I use as my workhorse model.

I don't know if my code is just "complex", but I find that Luna on max ignores the surrounding style and completely ignores logical consequences of a change, like just writing `del arg1, del arg2, ...` instead of dropping it from the surrounding code. All LLMs make questionable decisions at times, but Luna requires so much guidance that it's faster to just type it out yourself. What kind of routine tasks can one accomplish with such a model?
Post reply on HN