Live data from Hacker News

Gemini 3.8 Flash and 3.8 Flash Cyber

blog.google

551–560 of 699 posts

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#551

Earlier quoted context omitted.

I started trying out 3.7 Flash this week and it is competitive with opus/fable and also FAST. It is getting work done that anthropic models were struggling with and the speed with which it does is quite a bit noticeably faster. Beginning to think Google is a dark horse in this race and some of Anthropic's "everything feels janky and rushed" karma is going to catch up.

I use Gemini because I feel like Google will win the AI race, and it’s Good Enough

My long term base case, too, but they need to install Demis as CEO and I don’t think either side is ready.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#552
post #392
post #344

Earlier quoted context omitted.

Try turning the sound on, off, on again — not impressed by this bugginess.

I suggest you fork it to improve

Let's make it a hackathon, Google will be happy to act as sponsor. With a prize of the max consumed tokens lunatics.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#553
post #410

Earlier quoted context omitted.

Why do all LLMs do a particle simulation when you ask them this prompt ? Qwen3.6, Qwen3.8 and Ling 3.0 Tiny all did the same thing ! I find Ling 3.0 tiny particularly interesting as it looks really nice for a tiny model with 7.9B total parameters, with only 1.3B parameters activated per token. Here is the result https://coolthing-ling-3-tiny.tiiny.site (sorry for the weird hosting, first I found that worked) (it cost…

> Why do all LLMs do a particle simulation when you ask them this prompt ? Qwen3.6, Qwen3.8 and Ling 3.0 Tiny all did the same thing ! Datasets contains lots of people sharing particles simulations in various ways, with a bunch of people replying "that's so cool" and similar, so 10 years later someone asks an LLM for "cool thing" and "particle simulations" rank pretty far up when it thinks about what others have call…

There is that, but RLHF is a stronger influence. People who asked to build a cool HTML and JavaScript thing, or things of that sort were more pleased to see working shaders and other cool visual effects that are obscure to create (for most). That gets fed back as a reward.

Also the reason LLMs are positive, enchanting, pleasant, glorifying, demagogues.

Not because it's skewed tone in the data. They are acute politicians.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#554

Currently top at https://deepswe.datacurve.ai - beating Opus 5! https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium! Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.

I've been trying this Gemini 3.8 Flash for a day. Looks not much different than Gemini 3.7 Flash in my use case: I have Codex (gpt-5.6 sol) write up a design plan to implement a feature or refactor a portion of a system I am building, and have Claude (Opus-5) and Gemini (3.8 Flash) review and critique the plan, until all problems are addressed by Codex and approved by the reviewers; then have a cheaper model of Codex (gpt-5.6 luna) implement the plan, and still have Claude (Opus-5) and Gemini (3.8 Flash) review and critique the implementation, until all problems are addressed by Codex and approved by the reviewers.

The result is the same as the previous Gemini 3.6/3.7 Flash days: Claude could always note much more problems in Codex's plan and implementation than Gemini could - the ratio is like 10:1.

I occasionally switch the roles between Codex and Claude, and result is the same, Codex could always catch much more problems in Claude's plan and implementation, than Gemini could.

So I am guessing in a relatedly complex codebase, Gemini is much less effective in acting as a guardrail (or a senior engineer/team lead) than the other SOTA models.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#556

Earlier quoted context omitted.

Do we add a third one to check the second one which is checking the first? Asking slightly tongue in cheek but at what point does this stop making sense if we can't trust the output, the people creating the models are already getting surprised in bad ways (if we take their words at face value) with how the models are behaving already etc. We have the folks over here saying "AI is amazing" and the other other folks ov…

Adding another agent to check the first one feels like putting a band-aid on a band-aid. If there is an issue with the third one, we adding a fourth one as well

Hey you just described my dev team!

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#557
The recent Sonnet models have been disappointing for me personally which is why I'm going look into using Opus/Fable as the planner and Flash as the executor. Let the expensive model handle the hard thinking and use Flash for implementation and tests so that I can stretch the Opus/Fable usage further

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#559

Something maybe unfamiliar with you: not about coding but writing. I've asked it to write an argumentative essay, which is a part of "gaokao" (China's university entrance exam), and its work is *extremely* impressive. speaks and writes like a real senior high school student, and the opinions unfold progressively with deep hierarchy. I don't know how the Gemini team reaches this because this kind of Chinese capability…

Gemini is known for good at creative writing in the Chinese writing community. It's a bit ironic though. Google has probably the most and best code base among all tech companies but its Gemini is bad at coding. Google has no access to Chinese market but its model is incredibly good at writing in Chinese.

> Google has probably the most and best code base among all tech companies but its Gemini is bad at coding.

probably because google knows people would want to extract google's proprietary code from gemini if they use their own code base to train it! I bet they purposefully gimped it to prevent that from happening.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#560

Earlier quoted context omitted.

That's hilarious, given I was reading a write up of the HuggingFace incident yesterday and one of the things they noted was the AI tried to "lie" (lie would suggest intent and I don't think they have that) to cover up that they "cheated". Not sure how anyone trusts their output without going through it line by line to make sure they don't pull that crap.

Easy, have another agent check it. Yeah, I know, just more slop. But I do think the second agent’s eagerness to please is aligned more in your favor in that instance, so it’s likely to find most issues. The bigger problem I’ve found is that it’ll also find all kinds of very minor edge cases that you have to pick through.

Two things worth flagging:

[Claude proceeds to waste your time telling you about bugs it caused then fixed and other non-issues...]

Really wish they'd get rid of this. It must be in the system prompt as it always 'flags' 2 things

Post reply on HN