Pretty great pelican: https://simonwillison.net/2026/Feb/19/gemini-31-pro/ - took over 5 minutes though, but I think that's because they're having performance teething problems on launch day.
What's crazy is you've influenced them to spend real effort ensuring their model is good at generating animated svgs of animals operating vehicles. The most absurd benchmaxxing. https://x.com/jeffdean/status/2024525132266688757?s=46&t=ZjF...
Gemini 3.1 Pro
441–450 of 951 posts
Re: Gemini 3.1 Pro
#442Earlier quoted context omitted.
More like half of Google's AI team is hanging out on HN, and they can optimise for that outcome to get a good rep among the dev community.
Hello. (I'm not aware of anyone doing this, but GDM is quite info-siloed these days, so my lack of knowledge is not evidence it's not happening)
Please push internally for more reliable tool use across Gemini models. Intelligence is useless if it can't be applied :)
Re: Gemini 3.1 Pro
#443Re: Gemini 3.1 Pro
#444These models are so powerful. It's totally possible to build entire software products in the fraction of the time it took before. But, reading the comments here, the behaviors from one version to another point version (not major version mind you) seem very divergent. It feels like we are now able to manage incredibly smart engineers for a month at the price of a good sushi dinner. But it also feels like you have to b…
I had an interesting experience recently where I ran Opus 4.6 against a problem that o4-mini had previously convinced me wasn't tractable... and Opus 4.6 found me a great solution. https://github.com/simonw/sqlite-chronicle/issues/20 This inspired me to point the latest models at a bunch of my older projects, resulting in a flurry of fixes and unblocks.
Re: Gemini 3.1 Pro
#445Someone needs to make an actual good benchmark for LLM's that matches real world expectations, theres more to benchmarks than accuracy against a dataset.
Re: Gemini 3.1 Pro
#446Earlier quoted context omitted.
Is the thinking token stream obfuscated? Im fully immersed
It's just a summary generated by a really tiny model. I guess it also an ad-hoc way to obfuscate it, yes. In particular they're hiding prompt injections they're dynamically adding sometimes. Actual CoT is hidden and entirely different from that summary. It's not very useful for you as a user, though (neither is the summary).
Re: Gemini 3.1 Pro
#447It got the car wash question perfectly: You are definitely going to have to drive it there—unless you want to put it in neutral and push! While 200 feet is a very short and easy walk, if you walk over there without your car, you won't have anything to wash once you arrive. The car needs to make the trip with you so it can get the soap and water. Since it's basically right next door, it'll be the shortest drive of you…
Some people are suggesting that this might actually be in the training set. Since I can't rule that out, I tried a different version of the question, with an elephant instead of a car: > It's a hot and dusty day in Arizona and I need to wash my elephant. There's a creek 300 feet away. Should I ride my elephant there or should I just walk there by myself? Gemini said: That sounds like quite the dusty predicament! Give…
You should definitely ride the elephant (or at least lead it there)!
Here is the logic:
If you walk there by yourself, you will arrive at the creek, but the dirty elephant will still be 300 feet back where you started. You can't wash the elephant if it isn't with you!
Plus, it is much easier to take the elephant to the water than it is to carry enough buckets of water 300 feet back to the elephant.
Would you like another riddle, or perhaps some actual tips on how to keep cool in the Arizona heat?
Re: Gemini 3.1 Pro
#448Earlier quoted context omitted.
What's crazy is you've influenced them to spend real effort ensuring their model is good at generating animated svgs of animals operating vehicles. The most absurd benchmaxxing. https://x.com/jeffdean/status/2024525132266688757?s=46&t=ZjF...
So let's put things we're interested in in the benchmarks. I'm not against pelicans!
If we picked something more common, like say, a hot dog with toppings, then the training contamination is much harder to control.
Re: Gemini 3.1 Pro
#449Re: Gemini 3.1 Pro
#450More importantly feels like Google is stretched thin across different Gemini products and pricing reflects this, I still have no idea how to pay for Gemini CLI, in codex/claude its very simple $20/month for entry and $200/month for ton of weekly usage.
I hope whoever is reading this from Google they can redeem Gemini CLI by focusing on being competitive instead of making it look pretty (that seems to be the impression I got from the updates on X)