Live data from Hacker News

Gemini 3.1 Pro

blog.google

541–550 of 951 posts

Re: Gemini 3.1 Pro

#541
post #81
post #22

Earlier quoted context omitted.

Does the arc-agi-2 score more than doubling in a .1 release indicate benchmark-maxing? Though i dont know what arc-agi-2 actually tests

Theoretically, you can’t benchmaxx ARC-AGI, but I too am suspect of such a large improvement, especially since the improvement on other benchmarks is not of the same order.

https://arcprize.org/arc-agi/1/

It's a sort of arbitrary pattern matching thing that can't be trained on in the sense that the MMLU can be, but you can definitely generate billions of examples of this kind of task and train on it, and it will not make the model better on any other task. So in that sense, it absolutely can be.

I think it's been harder to solve because it's a visual puzzle, and we know how well today's vision encoders actually work https://arxiv.org/html/2407.06581v1

Re: Gemini 3.1 Pro

#542
post #312

It got the car wash question perfectly: You are definitely going to have to drive it there—unless you want to put it in neutral and push! While 200 feet is a very short and easy walk, if you walk over there without your car, you won't have anything to wash once you arrive. The car needs to make the trip with you so it can get the soap and water. Since it's basically right next door, it'll be the shortest drive of you…

The answer here is why I dislike Gemini, though it gets the correct answer, it's far too verbose.

I can't stand a model over-explaining, needless fluff and wasting tokens. I asked the question so I know the context.

Re: Gemini 3.1 Pro

#543

Gemini 3 is still in preview (limited rate limits) and 2.5 is deprecated (still live but won't be for long).[0] Are Google planning to put any of their models into production any time soon? Also somewhat funny that some models are deprecated without a suggested alternative(gemini-2.5-flash-lite). Do they suggest people switch to Claude? [0] https://ai.google.dev/gemini-api/docs/deprecations

I agree completely. I don't know how anyone can be building on these models when all of them are either deprecated or not actually released yet. As someone who has production systems running on the deprecated models, this situation really causes me grief.

I dont think any of them really wants api customers in the end. They are only temporarily useful.

Re: Gemini 3.1 Pro

#544
post #411

Earlier quoted context omitted.

This is part of the reason I don't like to use it. I feel it's hiding things from me, compared to other models that very clearly share what they are thinking.

To be fair, considering that the CoT exposed to users is a sanitized summary of the path traversal - one could argue that sanitized CoT is closer to hiding things than simply omitting it entirely.

This is something that bothers me. We had a beautiful trend on the Web of the browser also being the debugger - from View Source decades ago all the way up to the modern browser console inspired by Firebug. Everything was visible, under the hood, if you cared to look. Now, a lot of "thinking" is taking place under a shroud, and only so much of it can be expanded for visibility and insight into the process. Where is the option to see the entire prompt that my agent compiled and sent off, raw? Where's the option to see the output, replete with thinking blocks and other markup?

Re: Gemini 3.1 Pro

#545
post #156

This model says it accepts video inputs. I asked it to transcribe a 5 second video of a digital water curtain which spelled “Boo Happy Halloween”, and it came back with “Happy” which wasn’t the first frame, but also is incomplete. This kind of test is good because it requires stitching together info from the whole video.

It reads videos at 1fps by default. You have to set the video resolution to high in ai studio

This is inside the Gemini app.

Re: Gemini 3.1 Pro

#546

I hope this works better than 3.0 Pro I'm a former Googler and know some people near the team, so I mildly root for them to at least do well, but Gemini is consistently the most frustrating model I've used for development. It's stunningly good at reasoning, design, and generating the raw code, but it just falls over a lot when actually trying to get things done, especially compared to Claude Opus. Within VS Code Copi…

Gemini just doesn’t do even mildly well in agentic stuff and I don’t know why. OpenAI has mostly caught up with Claude in agentic stuff, but Google needs to be there and be there quickly

My guess is that Gemini team didn't focus on the large-scale RL training for the agentic workload. And they are trying to catch up with 3.1.

Re: Gemini 3.1 Pro

#547
post #518

Earlier quoted context omitted.

GPT-OSS-120b, a (downloadable) model released more than half a year ago also gets that right, I'm not sure this is such a great success. > Would you like me to check the local weather forecast to make sure it's not going to rain right after you wash it? Regardless of what I do, the days I decide to wash my car, it ALWAYS rains the day after, even if I look at the weather report and it says sunny. Cute that Gemini thi…

Non car person here. Why does that matter? It's not like rain means you didn't have to go to the wash, it rains often enough here that there wouldn't be car wash places left near me but there are plenty

Many people avoid washing cars just before rain to avoid spots, etc. Phoenix as an extreme example rarely rains and leaves everything filthy afterwards.

Re: Gemini 3.1 Pro

#548
People underrate Google's cost effectiveness so much. Half price of Opus. HALF.

Think about ANY other product and what you'd expect from the competition thats half the price. Yet people here act like Gemini is dead weight

____

Update:

3.1 was 40% of the cost to run AA index vs Opus Thinking AND SONNET, beat Opus, and still 30% faster for output speed.

https://artificialanalysis.ai/?speed=intelligence-vs-speed&m...

Re: Gemini 3.1 Pro

#549
post #17

Has anyone noticed that models are dropping ever faster, with pressure on companies to make incremental releases to claim the pole position, yet making strides on benchmarks? This is what recursive self-improvement with human support looks like.

Only using my historical experience and not Gemini 3.1 Pro, I think we see benchmark chasing then a grand release of a model that gets press attention... Then a few days later, the model/settings are degraded to save money. Then this gets repeated until the last day before the release of the new model. If we are benchmaxing this works well because its only being tested early on during the life cycle. By middle of the…

I have a relatively consistent task that it completed with new information on weekdays at the edge of its intelligence. Interestingly 3.0 flash was good when it came out, took a nose dive a month back and is now excellent, I actually can't fault it it's so good.

It's performance in antigravity has also actually improved since launch day where it was giving non-stop typescript errors (not sure if that was antigravity itself).

Re: Gemini 3.1 Pro

#550

It got the car wash question perfectly: You are definitely going to have to drive it there—unless you want to put it in neutral and push! While 200 feet is a very short and easy walk, if you walk over there without your car, you won't have anything to wash once you arrive. The car needs to make the trip with you so it can get the soap and water. Since it's basically right next door, it'll be the shortest drive of you…

Some people are suggesting that this might actually be in the training set. Since I can't rule that out, I tried a different version of the question, with an elephant instead of a car: > It's a hot and dusty day in Arizona and I need to wash my elephant. There's a creek 300 feet away. Should I ride my elephant there or should I just walk there by myself? Gemini said: That sounds like quite the dusty predicament! Give…

i would say this is a lower difficulty. the car question primes it to think about stuff like energy and pollution.
Post reply on HN