Google is terrible at marketing, but this feels like a big step forward. As per the announcement, Gemini 3.1 Pro score 68.5% on Terminal-Bench 2.0, which makes it the top performer on the Terminus 2 harness [1]. That harness is a "neutral agent scaffold," built by researchers at Terminal-Bench to compare different LLMs in the same standardized setup (same tools, prompts, etc.). It's also taken top model place on both…
Benchmarks aren't everything. Gemini consistently has the best benchmarks but the worst actual real-world results. Every time they announce the best benchmarks I try again at using their tools and products and each time I immediately go back to Claude and Codex models because Google is just so terrible at building actual products. They are good at research and benchmaxxing, but the day to day usage of the products an…
Gemini 3.1 Pro
131–140 of 951 posts
Re: Gemini 3.1 Pro
#132Pretty great pelican: https://simonwillison.net/2026/Feb/19/gemini-31-pro/ - took over 5 minutes though, but I think that's because they're having performance teething problems on launch day.
Re: Gemini 3.1 Pro
#133Every time I've used Gemini models for anything besides code or agentic work they lean so far into the RLHF induced bold lettering and bullet point list barf that everything they output reads as if the model was talking _at_ me and not _with_ me. In my Openclaw experiment(s) and in the Gemini web UI, I've specifically added instructions to avoid this type of behavior, but it only seemed to obey those rules when I rem…
If a model doesn't optimize the formatting of its output display for readability, I don't want to read it.
Tables, embedded images, use of bulleted lists and bold/italicizing etc.
Re: Gemini 3.1 Pro
#134I hope to have great next two weeks before it gets nerfed.
I've found Google (at least in AI Studio) are the only provider NOT to nerf their models after a few weeks
Re: Gemini 3.1 Pro
#135ok , so they are scared that 5.3 (pro) will be released today/tomorrow and blow it out of the water and rushed it while they could still reference 5.2 benchmarks.
I don't think models blow other models anymore. We have the big 3 which are neck to neck in most benchmarks and the rest. I doubt that 5.3 will blow the others.
Re: Gemini 3.1 Pro
#136This kind of test is good because it requires stitching together info from the whole video.
Re: Gemini 3.1 Pro
#137Earlier quoted context omitted.
It's an excellent demonstration of the main issue I have with the Gemini family of models, they always go "above and beyond" to do a lot of stuff, even if I explicitly prompt against it. In this case, most of the SVG ends up consisting not just of a bike and a pelican, but clouds, a sun, a hat on the pelican and so much more. Exactly the same thing happens when you code, it's almost impossible to get Gemini to not do…
Do you have Personalization Instructions set up for your LLM models? You can make their responses fairly dry/brief.
Re: Gemini 3.1 Pro
#138I always try Gemini models when they get updated with their flashy new benchmark scores, but always end up using Claude and Codex again... I get the impression that Google is focusing on benchmarks but without assessing whether the models are actually improving in practical use-cases. I.e. they are benchmaxing Gemini is "in theory" smart, but in practice is much, much worse than Claude and Codex.
Which cases? Not trying to sound bad but you didn't even provide of cases you are using Claude\Codex\Gemini for.
Re: Gemini 3.1 Pro
#139Gemini 3 is pretty good, even Flash is very smart for certain things, and fast! BUT it is not good at all at tool calling and agentic workflows, especially compared to the recent two mini-generations of models (Codex 5.2/5.3, the last two versions of Anthropic models), and also fell behind a bit in reasoning. I hope they manage to improve things on that front, because then Flash would be great for many tasks.