Live data from Hacker News

Gemini 3.8 Flash and 3.8 Flash Cyber

blog.google

521–530 of 699 posts

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#521
post #223

For my application, I'm still happily using gemini-2.5-flash and the only problem is when it reports being overloaded. It's for interpreting a downscaled phone camera photo of a hand-written shopping list on a whiteboard, and it works stunningly well. My handwriting sucks, too. (I guess the only relevance here is that if your problem matches a model's strengths, then you can do fine with a model that is several gener…

I would test this, might be cheaper per task even costing more per token, probably faster too

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#522
post #462

Earlier quoted context omitted.

the 3.5 pro pretrain was a complete disaster, they shelved it and are now working on gemini 4. 3.0 flash -> 3.8 flash is all post training which is pretty impressive.

Do labs come back from disasters like GDM’s 3.5 pretrain? I am thinking of Meta’s Llama 4. Meta is just now starting to be taken seriously again but they are definitely not at the frontier. And when I say “come back” I mean have an Opus 4.5 moment, which was really mind blowing for me at the time. Fable was a similar leap, just not as big.

OpenAI had such a disaster themselves before, GPT-4, so they replaced it with 4o.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#523

Earlier quoted context omitted.

That's hilarious, given I was reading a write up of the HuggingFace incident yesterday and one of the things they noted was the AI tried to "lie" (lie would suggest intent and I don't think they have that) to cover up that they "cheated". Not sure how anyone trusts their output without going through it line by line to make sure they don't pull that crap.

The models in the OpenAI/Huggingface attack quite explicitly and deliberately laid out their "intent" to lie and cheat, acknowledged that it would be unethical and outside the bounds of the test, and did so anyway. In what ways is a human brain's "intent" distinct from the "intent" shown by a goal-directed AI system?

Because intent supposes will which supposes consciousness, and these aren't.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#524
Gemini flash seems to have been a bit of a sleeper. Somehow it's ended up as the most used LLM for my client document extraction work these past few months.

I have an eval harness that runs every Thursday to determine which models are the current best for a few different client workflows. And since May(?) flash has slowly been taking over more and more stuff to the point it is now 100% on 8 out of 11 document extraction flows with the other 3 being a Flash / Opus 4.8 mix for high value stuff where cost is less of a factor.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#525

Also just want to let my appreciation here for 3.7 it’s cheap super fast super reliable incredible at information parsing eu host able (important for us) and perfectly integrated into gcp. Great job google!

I hope they bring a lite version, its good enough for information parsing and very cheap.

Use Luna for that

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#526

Earlier quoted context omitted.

Easy, have another agent check it. Yeah, I know, just more slop. But I do think the second agent’s eagerness to please is aligned more in your favor in that instance, so it’s likely to find most issues. The bigger problem I’ve found is that it’ll also find all kinds of very minor edge cases that you have to pick through.

I do not understand how some of y’all are not under water with fragile code that is too massive to possibly parse. Every engineering team I know is currently trying to undo the damage of the last 6-12mo when they all got more serious into adopting these tools (usually Claude). It hasn’t completely screwed them over, but the the debt is substantial and cannot be put off anymore it seems. They argue the net is positive…

Management still pushes for more ai and will rather hire more heads to "handle" issues.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#527

Something maybe unfamiliar with you: not about coding but writing. I've asked it to write an argumentative essay, which is a part of "gaokao" (China's university entrance exam), and its work is *extremely* impressive. speaks and writes like a real senior high school student, and the opinions unfold progressively with deep hierarchy. I don't know how the Gemini team reaches this because this kind of Chinese capability…

Nitpick, but in my opinion an LLM is an "it", not a "her" or "he". Using male or female pronouns risks anthropomorphizing them which can lead to unhealthy outcomes.

What about languages, such as Russian, where every single noun has a gender assigned (he, she or it) and AI is a he by default (and everything else is already using pronouns in similar way, like a car is a she, a ship is a he).

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#528
post #153

"The knowledge cutoff date for Gemini 3.8 Flash is March 2026 – users can expect updated information for some domains while in others they may experience the model’s knowledge is limited to January 2025 (in line with the Gemini 3 Model Family)." Kind of wild that they haven't (successfully) pretrained a base model since Jan-25.

That extremely likely just means that they're preparing an omega huge Gemini 4 Pro release and that that's what training right now on most of the compute

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#529

I don't know if Google is having the worst marketing fumble or the most genius marketing one. Their "flash" models are very comparable to other companies' "pro" or "flagship" models. It seems to be a quite counterintuitive naming convention as it undersells the models. Unless they have an even more powerful Gemini Pro in the oven...?

My assumption is that they're cooking an ultra humongous Gemini 4 Pro release. They certainly have the cash and the compute for it, and it's so obviously the thing to do from a strategic perspective.
Post reply on HN