Live data from Hacker News

Gemini 3.8 Flash and 3.8 Flash Cyber

blog.google

341–350 of 699 posts

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#341
How generous is the Google subscription quotas compared to Anthropic and OpenAI? This sounds like a really good potential model for high volume due to its speed and cost effectiveness.

(By high volume I mean things like "main app just updated with XYZ commits, please scan XYZ plugins and surface any compatibility issues")

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#342
post #85

Earlier quoted context omitted.

The Gemini models have openly trained for SVG output, apparently with a specialism on animals in forms of transport! https://twitter.com/JeffDean/status/2024525132266688757

Community effort happening here to build the ideal dataset: https://github.com/scosman/pelicans_riding_bicycles

Wow nice. If an llm could replicate these excellent examples, I'd consider the pelican benchmark fully saturated.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#343
post #200

The speed combined with the fact that this thing is really good at HTML JavaScript is pretty exciting. Here's what I got for 1.8 cents and 13 seconds from the prompt "make me a cool thing in html": https://gisthost.github.io/?6a77bc41a81718c6aaa10d4ab243c59f Transcript here (it was part of a chat): https://gist.github.com/simonw/b6149a49d327164d67d62c3d12992...

Definitely cool. I noticed it felt a little janky on my PC despite being "60 FPS"...then I noticed the "60 FPS" is hard-coded into the HTML.

That's hilarious, given I was reading a write up of the HuggingFace incident yesterday and one of the things they noted was the AI tried to "lie" (lie would suggest intent and I don't think they have that) to cover up that they "cheated".

Not sure how anyone trusts their output without going through it line by line to make sure they don't pull that crap.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#344
post #200

The speed combined with the fact that this thing is really good at HTML JavaScript is pretty exciting. Here's what I got for 1.8 cents and 13 seconds from the prompt "make me a cool thing in html": https://gisthost.github.io/?6a77bc41a81718c6aaa10d4ab243c59f Transcript here (it was part of a chat): https://gist.github.com/simonw/b6149a49d327164d67d62c3d12992...

Definitely cool. I noticed it felt a little janky on my PC despite being "60 FPS"...then I noticed the "60 FPS" is hard-coded into the HTML.

Try turning the sound on, off, on again — not impressed by this bugginess.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#345

Earlier quoted context omitted.

Same here. I see so many people obsessing over the latest most state of the art bleeding edge models and yelling at Google for not being there, but I feel like the vast majority of people don't actually need those models. Flash has just been super useful and incredibly fast in my experience.

I prefer luna for most development, especially when I am guiding the process. Sometimes terra. I have had terrible results coding with sol. It is way over-tuned on RL to make something that completes the task, no matter what. I end up with way too much code that does a lot of things I didn't ask for.

I setup Luna as main Claude Code driver (so zero anthropic api use) and it nailed crisply a handful of python tasks, gonna continue this way.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#347
post #162

Earlier quoted context omitted.

They are all much larger and more expensive models. Google does not have a frontier model right now, but for cheap ones, they are better than event the chinese models now.

That's not being debated here. The initial reported numbers were false and this was simply pointed out. You're changing the subject.

> [...] shows an intelligence score of 59, the same as Opus 5 medium!

Nothing here is false, you are simply confused. You either didn't read what they wrote in its entirety or decided to reinterpret what they did write.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#348
post #314
post #309

People have been sleeping on Gemini lately but these last few Flash releases (which were very rapid) are damn good. These sort of fast and cheap models are great for tasks that are verifiable and can be retried infinitely (like coding), you can basically get frontier results with a good harness (at a fraction of the time and money).

As someone who has stubbornly stuck with Claude Code, what's a good harness for Gemini models?

Antigravity is probably the best of the bunch I've tried. I'd say it's pretty comparable to Claude Code (I use both daily).

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#349

Earlier quoted context omitted.

Google One plans are quite a good value actually - for a few bucks you get more Gemini plus space in Drive and other extras. Even through API, $3.75 for nearly Sol-level quality isn't that bad. And let's not forget you can use it for free in AI Studio, and in the user app (even free accounts get tons of usage, though it's still 3.6 there), and in Antygravity.

That's the thing. I am completely lost because there are so many redundant paths to get the same thing and I'm trying to figure out which one is the best deal

Just put Mythos on the task; it’ll work out the best way in a measly few hours.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#350
post #207
post #41

Earlier quoted context omitted.

Crushing it on DeepSWE is a very big deal. Excited to give this a try.

> DeepSWE is a very big deal It's clearly been "dealt with" already. When it launched we had interesting gaps and definitely differences. Now every new release is "crushing it".

Will look forward to the "feel" of the model in real testing. But I agree that these benchmarks do get "dealt with" rapidly. That's a shame, but I guess it's the times we live in.
Post reply on HN