Live data from Hacker News

Gemini 3.1 Pro

blog.google

491–500 of 951 posts

Re: Gemini 3.1 Pro

#491
post #451

Earlier quoted context omitted.

What's crazy is you've influenced them to spend real effort ensuring their model is good at generating animated svgs of animals operating vehicles. The most absurd benchmaxxing. https://x.com/jeffdean/status/2024525132266688757?s=46&t=ZjF...

I like how they also did a frog on a penny-farthing and a giraffe driving a tiny car and an ostrich on roller skates and a turtle kickflipping a skateboard and a dachshund driving a stretch limousine.

reminds me of andor, luthen, positive reinforcing wasting time of emperor

Re: Gemini 3.1 Pro

#494
post #486

Earlier quoted context omitted.

Remember when ARC 1 was basically solved, and then ARC 2 (which is even easier for humans) came out, and all of the sudden the same models that were doing well on ARC 1 couldn’t even get 5% on ARC 2? Not convinced these benchmark improvements aren’t data leakage.

ARC 2 was made specifically to artificially lower contemporary LLM scores, therefore any kind of model improvements will have outsized effects Also people use "saturated" too liberally. The top left corner 1 cent per task is saturated IMO. Since there are billions of people who would perfer to solve arc 1 tasks at 52 cents per task. Arc 2 a human would make thousands of dollars a day with 99.99% accuracy

You are saying something interesting but too esoteric. Can you explain for beginners?

Re: Gemini 3.1 Pro

#495
post #52

Pretty great pelican: https://simonwillison.net/2026/Feb/19/gemini-31-pro/ - took over 5 minutes though, but I think that's because they're having performance teething problems on launch day.

Cost per task has increased 4.2x but their ARC-AGI-2 score went from 33.6% to 77.1%

Cost per task is still significantly lower than Opus. Even Opus 4.5

https://arcprize.org/leaderboard

Re: Gemini 3.1 Pro

#496

Earlier quoted context omitted.

Good to see it wearing a helmet. Their safety team must be on their game.

Yes but why would a pelican need a helmet? If it falls over it can just fly away... Common sense 1 Gemini 0

Obviously these domestic pelicans can't fly, otherwise why would they need a bike?

Re: Gemini 3.1 Pro

#498

Earlier quoted context omitted.

A few thoughts: - One thing to be aware of is that LLMs can be much smarter than their ability to articulate that intelligence in words. For example, GPT-3.5 Turbo was beastly at chess (1800 elo?) when prompted to complete PGN transcripts, but if you asked it questions in chat, its knowledge was abysmal. LLMs don't generalize as well as humans, and sometimes they can have the ability to do tasks without the ability t…

We’re literally at the point where trillions of dollars have been invested in these things and the surrounding harnesses and architecture, and they still can’t do economically useful work on their own. You’re way too bullish here.

Neither do cars until very recently. A tool doesn't have to be unsupervised to be useful.

Re: Gemini 3.1 Pro

#499
post #376

Gets 10/10 on my potato benchmarks: https://aibenchy.com/model/google-gemini-3-1-pro-preview-med...

Are you intentionally keeping the benchmarks private?

Yes.

I am trying to think what's the best way to give most information about how the AI models fail, without revealing information that can help them overfit on those specific tests.

I am planning to add some extra LLM calls, to summarize the failure reason, without revealing the test.

Re: Gemini 3.1 Pro

#500

Gemini 3 is still in preview (limited rate limits) and 2.5 is deprecated (still live but won't be for long).[0] Are Google planning to put any of their models into production any time soon? Also somewhat funny that some models are deprecated without a suggested alternative(gemini-2.5-flash-lite). Do they suggest people switch to Claude? [0] https://ai.google.dev/gemini-api/docs/deprecations

Have 2.5 in prod. Hope they release 3 lite soon so it will be easier to swap them. Holding my breath as pro pricing is a non starter.
Post reply on HN