Gemini 3.1 Pro
21–30 of 951 posts
Re: Gemini 3.1 Pro
#22blog post is up- https://blog.google/innovation-and-ai/models-and-research/ge... edit: biggest benchmark changes from 3 pro: arc-agi-2 score went from 31.1% -> 77.1% apex-agents score went from 18.4% -> 33.5%
Re: Gemini 3.1 Pro
#23Knowledge cutoff is unchanged at Jan 2025. Gemini 3.1 Pro supports "medium" thinking where Gemini 3 did not: https://ai.google.dev/gemini-api/docs/gemini-3
Compare to Opus 4.6's $5/M input, $25/M output. If Gemini 3.1 Pro does indeed have similar performance, the price difference is notable.
Re: Gemini 3.1 Pro
#24Re: Gemini 3.1 Pro
#25It's only February...
Re: Gemini 3.1 Pro
#26Re: Gemini 3.1 Pro
#27Re: Gemini 3.1 Pro
#28Has anyone noticed that models are dropping ever faster, with pressure on companies to make incremental releases to claim the pole position, yet making strides on benchmarks? This is what recursive self-improvement with human support looks like.
Then a few days later, the model/settings are degraded to save money. Then this gets repeated until the last day before the release of the new model.
If we are benchmaxing this works well because its only being tested early on during the life cycle. By middle of the cycle, people are testing other models. By the end, people are not testing them, and if they did it would barely shake the last months of data.
Re: Gemini 3.1 Pro
#29Has anyone noticed that models are dropping ever faster, with pressure on companies to make incremental releases to claim the pole position, yet making strides on benchmarks? This is what recursive self-improvement with human support looks like.