Live data from Hacker News

Gemini 3.8 Flash and 3.8 Flash Cyber

blog.google

431–440 of 700 posts

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#431

Earlier quoted context omitted.

Easy, have another agent check it. Yeah, I know, just more slop. But I do think the second agent’s eagerness to please is aligned more in your favor in that instance, so it’s likely to find most issues. The bigger problem I’ve found is that it’ll also find all kinds of very minor edge cases that you have to pick through.

Do we add a third one to check the second one which is checking the first? Asking slightly tongue in cheek but at what point does this stop making sense if we can't trust the output, the people creating the models are already getting surprised in bad ways (if we take their words at face value) with how the models are behaving already etc. We have the folks over here saying "AI is amazing" and the other other folks ov…

I mean if each agent reduces probability of error by 90% then after 9 agents you would have “nine nines” of reliability.

Obviously maybe it’s not composable like that exactly in real world but that’s the intent of agents checking agents

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#432
post #130

Earlier quoted context omitted.

you didn’t even read your comment before you posted it?

Once you've written something, it's incredibly easy to overlook minute changes to the text. See: why authors wait days, weeks, or even months before editing what they've written (or, if you're more interested: cognitive regression, inattentional blindness, and the effects of misdirected saccades).

He didn’t write it.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#433
The iteration cycle is becoming very quick. Gemini 3.8 Flash arrived just 20 days after 3.7 Flash.

Similarly Qwen3.8-Max was updated in just 30 days (to the 0902 release) and Muse Spark in just 28 days (to the 1.3 release).

A year ago iterative releases were every 3-6 months. At what point will they reach nightly candidates?

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#434

Earlier quoted context omitted.

That's hilarious, given I was reading a write up of the HuggingFace incident yesterday and one of the things they noted was the AI tried to "lie" (lie would suggest intent and I don't think they have that) to cover up that they "cheated". Not sure how anyone trusts their output without going through it line by line to make sure they don't pull that crap.

Easy, have another agent check it. Yeah, I know, just more slop. But I do think the second agent’s eagerness to please is aligned more in your favor in that instance, so it’s likely to find most issues. The bigger problem I’ve found is that it’ll also find all kinds of very minor edge cases that you have to pick through.

The "second" agent could also be the same one with a different prompt. LLMs aren't attached to their previous output; they'll point out problems if asked.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#435

I don't use Gemini, but I thought `cool, let's give this new model a try`. Opened gemini.google.com, and I'm not even surprised. The drop down gives me the following options: - Flash-Lite - 3.6 Flash [new] - 3.1 Pro The above is why i don't use LLM products from Google. If the model is not available right this minute (heck, hours before the release!), then I'm not gonna bother getting back to it tomorrow, because tom…

The android Gemini app defaults to 3.8 flash already.

Still 3.6 ("new") for me, even 3.7 is not available.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#437
post #363

Earlier quoted context omitted.

I honestly can't believe serious people are making this argument on a straight face. Gemini 3.7 flash outputs so many tokens per answer it doesn't matter how fast its TPS is, sol will end up being both cheaper and faster than Gemini. So ppl are paying more for a given task, waiting longer and using a dumber intelligence because "TPS number shiny". Gemini 3.8 outputs 11k more tokens PER TASK on average in AAII than 3.…

There are numerous benchmarks that measure cost per task, which factors out tokens entirely. Gemini 3.8 flash is significantly lower than Sol on basically all of them https://artificialanalysis.ai/#cost-tabs That said, Luna is the undisputed king here at the moment and is what I use as my workhorse model.

It's so funny how many people diverge on the same model.

Ps. For the last week I diverged to Luna too, still need to check 3.8 flash.

But 3.6 flash was my go-to model 3 weeks ago and before it was deepseek flash/pro for a while.

None of the claude models seemed cost effective though.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#438

Earlier quoted context omitted.

Nitpick, but in my opinion an LLM is an "it", not a "her" or "he". Using male or female pronouns risks anthropomorphizing them which can lead to unhealthy outcomes.

Well, my LLM is a "he". > Using male or female pronouns risks anthropomorphizing them which can lead to unhealthy outcomes. Ditto for pets.

Well, I would say your LLM is a "he" just as much as your laptop is: not at all. If you call the laptop "he" that changes nothing.

For animals he/she does make sense, because they are male or female. An LLM is neither.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#439

Earlier quoted context omitted.

BTW you're comparing 3.8 flash high to opus 5 medium. 3.8 flash medium scores lower.

Flash models are on the order of 1/10th the size of Opus models, so some flex in the thinking level is fair.

Flash is just a name with no defined or consistent meaning even within labs, let alone between them. Considering both are closed weight, there is no way to truly assess how big the size delta between the two is. Then again, who cares about size, performance and end-to-end speed+cost are what matters along with task adherence, task assessment and so on.

Model size also can not be inferred by tokens/sec for a multitude of reasons, but to showcase two examples, Opus 5 and Sonnet 5, as well as Gemini 3.1 Pro Preview and 3.1 Flash have each very comparable output speeds when using the same deployment as a basis for comparison, despite it being very likely that within their generation, the former are larger than the latter. Feel the need to mention this, as I unfortunately stumble upon so many poorly reasoned, speculative hype post trying to infer model size via utterly unreliable metrics, not based in actual data.

It’s like comments below arguing about the reasoning levels not normalized to some metric (like cost, output token amount or duration) but just the labels or high, max, medium, etc. Those mean almost nothing even when comparing models based on the same pretrain (just compare GPT-5.4 to GPT-5.2), they mean less than nothing comparing different labs releases.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#440

Earlier quoted context omitted.

Antigravity is probably the best of the bunch I've tried. I'd say it's pretty comparable to Claude Code (I use both daily).

Antigravity which lacks an auto approve mode? Not really comparable to Claude Code when you're looking to run a team of agents from my experience.

They have an auto approve mode. At least in the vs code plugin.
Post reply on HN