Earlier quoted context omitted.
Weren't we barely scraping 1-10% on this with state of the art models a year ago and it was considered that this is the final boss, ie solve this and its almost AGI-like? I ask because I cannot distinguish all the benchmarks by heart.
Yes, but benchmarks like this are often flawed because leading model labs frequently participate in 'benchmarkmaxxing' - ie improvements on ARC-AGI2 don't necessarily indicate similar improvements in other areas (though it does seem like this is a step function increase in intelligence for the Gemini line of models)
Gemini 3 Deep Think
81–90 of 722 posts
Re: Gemini 3 Deep Think
#82Earlier quoted context omitted.
If you look at the problem space it is easy to see why it's toast, maybe there's intelligence in there, but hardly general.
the best way I've seen this describes is "spikey" intelligence, really good at some points, those make the spikes humans are the same way, we all have a unique spike pattern, interests and talents ai are effectively the same spikes across instances, if simplified. I could argue self driving vs chatbots vs world models vs game playing might constitute enough variation. I would not say the same of Gemini vs Claude vs .…
So maybe we are forced to be more balanced and general whereas AI don't have to.
Re: Gemini 3 Deep Think
#83The pelican riding a bicycle is excellent . I think it's the best I've seen. https://simonwillison.net/2026/Feb/12/gemini-3-deep-think/
Highly disagree. I was expecting something more realistic... the true test of what you are doing is how representative is the thing in relation to the real world. E.g. does the pelican look like a pelican as it exists in reality? This cartoon stuff is cute but doesnt pass muster in my view. If it doesn't relate to the real world, then it most likely will have no real effect on the real economy. Pure and simple.
Re: Gemini 3 Deep Think
#84Less than a year to destroy Arc-AGI-2 - wow.
Re: Gemini 3 Deep Think
#85Less than a year to destroy Arc-AGI-2 - wow.
It's a useless meaningless benchmark though, it just got a catchy name, as in, if the models solve this it means they have "AGI", which is clearly rubbish. Arc-AGI score isn't correlated with anything useful.
Re: Gemini 3 Deep Think
#86Earlier quoted context omitted.
https://arcprize.org/leaderboard $13.62 per task - so we need another 5-10 years for the price to run this to become reasonable? But the real question is if they just fit the model to the benchmark.
That's not a long time in the grand scheme of things.
Re: Gemini 3 Deep Think
#87Earlier quoted context omitted.
the best way I've seen this describes is "spikey" intelligence, really good at some points, those make the spikes humans are the same way, we all have a unique spike pattern, interests and talents ai are effectively the same spikes across instances, if simplified. I could argue self driving vs chatbots vs world models vs game playing might constitute enough variation. I would not say the same of Gemini vs Claude vs .…
You can get more spiky with AIs, whereas with human brain we are more hard wired. So maybe we are forced to be more balanced and general whereas AI don't have to.
Why is it so easy for me to open the car door, get in, close the door, buckle up. You can do this in the dark and without looking.
There are an infinite number of little things like this you think zero about, take near zero energy, yet which are extremely hard for Ai
Re: Gemini 3 Deep Think
#88Re: Gemini 3 Deep Think
#89Earlier quoted context omitted.
Trick? Lol not a chance. Alphabet is a pure play tech firm that has to produce products to make the tech accessible. They really lack in the latter and this is visible when you see the interactions of their VP's. Luckily for them, if you start to create enough of a lead with the tech, you get many chances to sort out the product stuff.
You sound like Russ Hanneman from SV