Earlier quoted context omitted.
Weren't we barely scraping 1-10% on this with state of the art models a year ago and it was considered that this is the final boss, ie solve this and its almost AGI-like? I ask because I cannot distinguish all the benchmarks by heart.
Here's a good thread over 1+ month, as each model comes out https://bsky.app/profile/pekka.bsky.social/post/3meokmizvt22... tl;dr - Pekka says Arc-AGI-2 is now toast as a benchmark
Gemini 3 Deep Think
51–60 of 722 posts
Re: Gemini 3 Deep Think
#52The pelican riding a bicycle is excellent . I think it's the best I've seen. https://simonwillison.net/2026/Feb/12/gemini-3-deep-think/
How likely this problem is already on the training set by now?
Re: Gemini 3 Deep Think
#53Arc-AGI-2: 84.6% (vs 68.8% for Opus 4.6) Wow. https://blog.google/innovation-and-ai/models-and-research/ge...
$13.62 per task - so we need another 5-10 years for the price to run this to become reasonable?
But the real question is if they just fit the model to the benchmark.
Re: Gemini 3 Deep Think
#54Re: Gemini 3 Deep Think
#55I'm pretty certain that DeepMind (and all other labs) will try their frontier (and even private) models on First Proof [1]. And I wonder how Gemini Deep Think will fare. My guess is that it will get half the way on some problems. But we will have to take an absence as a failure, because nobody wants to publish a negative result, even though it's so important for scientific research. [1] https://1stproof.org/
Re: Gemini 3 Deep Think
#56Re: Gemini 3 Deep Think
#57Arc-AGI-2: 84.6% (vs 68.8% for Opus 4.6) Wow. https://blog.google/innovation-and-ai/models-and-research/ge...
It's completely misnamed. It should be called useless visual puzzle benchmark 2.
It's a visual puzzle, making it way easier for humans than for models trained on text firstly. Secondly, it's not really that obvious or easy for humans to solve themselves!
So the idea that if an AI can solve "Arc-AGI" or "Arc-AGI-2" it's super smart or even "AGI" is frankly ridiculous. It's a puzzle that means nothing basically, other than the models can now solve "Arc-AGI"
Re: Gemini 3 Deep Think
#58I need to test the sketch creation a s a p. I need this in my life because learning to use Freecad is too difficult for a busy person like me (and frankly, also quite lazy)
Re: Gemini 3 Deep Think
#59The pelican riding a bicycle is excellent . I think it's the best I've seen. https://simonwillison.net/2026/Feb/12/gemini-3-deep-think/
It was sort of humorous for the maybe first 2 iterations, now it's tacky, cheesy, and just relentless self-promotion.
Again, like I said before, it's also a terrible benchmark.