Earlier quoted context omitted.
Models are soon going to start benchmaxxing generating SVGs of pelicans on bikes
Soon? I'd be willing to bet it's been included in the training set at least 6 months by now. Not so obvious so it generates always perfect pelicans on bikes, but sufficiently for the "minibench" to be less useful today than in the past.
Gemini 3.1 Pro
501–510 of 951 posts
Re: Gemini 3.1 Pro
#502Earlier quoted context omitted.
Is the thinking token stream obfuscated? Im fully immersed
It's just a summary generated by a really tiny model. I guess it also an ad-hoc way to obfuscate it, yes. In particular they're hiding prompt injections they're dynamically adding sometimes. Actual CoT is hidden and entirely different from that summary. It's not very useful for you as a user, though (neither is the summary).
Re: Gemini 3.1 Pro
#503Re: Gemini 3.1 Pro
#504Gets 10/10 on my potato benchmarks: https://aibenchy.com/model/google-gemini-3-1-pro-preview-med...
Re: Gemini 3.1 Pro
#505Can we switch from Claude Code to Google yet? Benchmarks are saying: just try But real world could be different
My sense is that the Gemini models are very capable but the Gemini CLI experience is subpar compared to Claude Code and Codex. I'm guess that it's the harness but since it can get confused, fall into doom loops, and generally lose the plot in a way that the model does not in Gemini Studio or the Gemini app. I think a bunch of these harnesses are open source so it surprises me that there can be such a gulf between the…
It goes into loops and never completes a task 8 times out of 10 that i've used it.
Re: Gemini 3.1 Pro
#506Why don't they show Grok benchmarks?
Re: Gemini 3.1 Pro
#507Re: Gemini 3.1 Pro
#508Re: Gemini 3.1 Pro
#509Earlier quoted context omitted.
What's crazy is you've influenced them to spend real effort ensuring their model is good at generating animated svgs of animals operating vehicles. The most absurd benchmaxxing. https://x.com/jeffdean/status/2024525132266688757?s=46&t=ZjF...
Animated SVG is huge. People in different professions are worrying to different degrees in terms of being replaced by ML, but this one is huge with regards to digital art.
I've been meaning to let coding agents take a stab at using the lottie library https://github.com/airbnb/lottie-web to supercharge the user experience without needing to make it a full time job
Re: Gemini 3.1 Pro
#510Earlier quoted context omitted.
your question may have become part of the training data with how much coverage there was around it. perhaps you should devise a new test :P
3.1 Pro has the same Jan 2025 knowledge cutoff as the other 3 series models. So if 3.1 has it in its training data, the other ones would have as well.