Earlier quoted context omitted.
I look forward to them trying. I'll know when the pelican riding a bicycle is good but the ocelot riding a skateboard sucks.
Would it not be better to have 100 such tests "Pelican on bicycle", "Tiger on stilts"..., and generate them all for every new model but only release a new one each time. That way you could show progression across all models, attempts at benchmaxxing would be more obvious. Given the crazy money and vying for supremacy among AI companies right now it does seem naive to belive that no attempt at better pelicans on bicyc…
Gemini 3 Deep Think
641–650 of 722 posts
Re: Gemini 3 Deep Think
#642The pelican riding a bicycle is excellent . I think it's the best I've seen. https://simonwillison.net/2026/Feb/12/gemini-3-deep-think/
Do you have to still keep trying to bang on about this relentlessly? It was sort of humorous for the maybe first 2 iterations, now it's tacky, cheesy, and just relentless self-promotion. Again, like I said before, it's also a terrible benchmark.
Re: Gemini 3 Deep Think
#643Earlier quoted context omitted.
Wartime Google gave us Google+. Wartime Google is still bumbling, and despite OpenAI's numerous missteps, I don't think it has to worry about Google hurting its business yet.
Google+ was fun. Failed in the market though. Apple made a social network called Ping. Disaster. MobileMe was silly. Microsoft made Zune and the Kin 1 and Kin 2 devices and Windows phone and all sorts of other disasters. These things happen.
Re: Gemini 3 Deep Think
#644Earlier quoted context omitted.
That's not really how it works, the recent Erdos proofs in Lean were done by a specialized proprietary model (Aristotle by Harmonic) that's specifically trained for this task. Normal agents are not effective.
Why did you omit the other AI-generated Erdos proofs not done by a proprietary model, which occurred on timescales stretched across significantly longer time than 5 days?
Re: Gemini 3 Deep Think
#645Earlier quoted context omitted.
> If something claims to be conscious and we can't prove it's not, we have no choice but to believe it. This is not a good test. A dog won't claim to be conscious but clearly is, despite you not being able to prove one way or the other. GPT-3 will claim to be conscious and (probably) isn't, despite you not being able to prove one way or the other.
An LLM will claim whatever you tell it to claim. (In fact this Hacker News comment is also conscious.) A dog won’t even claim to be a good boy.
Last week gemini argued with me about an auxiliary electrical generator install method and it turned out to be right, even though I pushed back hard on it being incorrect. First time that has ever happened.
Re: Gemini 3 Deep Think
#646Earlier quoted context omitted.
>Why is it so easy for me to open the car door Because this part of your brain has been optimized for hundreds of millions of years. It's been around a long ass time and takes an amazingly low amount of energy to do these things. On the other hand the 'thinking' part of your brain, that is your higher intelligence is very new to evolution. It's expensive to run. It's problematic when giving birth. It's really slow wi…
> There's a term for this, but I can't think of it at the moment. Moravec's paradox: https://epoch.ai/gradient-updates/moravec-s-paradox
Re: Gemini 3 Deep Think
#647Earlier quoted context omitted.
François Chollet, creator of ARC-AGI, has consistently said that solving the benchmark does not mean we have AGI. It has always been meant as a stepping stone to encourage progress in the correct direction rather than as an indicator of reaching the destination. That's why he is working on ARC-AGI-3 (to be released in a few weeks) and ARC-AGI-4. His definition of reaching AGI, as I understand it, is when it becomes i…
> His definition of reaching AGI, as I understand it, is when it becomes impossible to construct the next version of ARC-AGI because we can no longer find tasks that are feasible for normal humans but unsolved by AI. That is the best definition I've yet to read. If something claims to be conscious and we can't prove it's not, we have no choice but to believe it. Thats said, I'm reminded of the impossible voting tests…
Re: Gemini 3 Deep Think
#648Earlier quoted context omitted.
The idea that an AI lab would pay a small army of human artists to create training data for $animal on $transport just to cheat on my stupid benchmark delights me.
When you're spending trillions on capex, paying a couple of people to make some doodles in SVGs would not be a big expense.
Re: Gemini 3 Deep Think
#649Earlier quoted context omitted.
Weren't we barely scraping 1-10% on this with state of the art models a year ago and it was considered that this is the final boss, ie solve this and its almost AGI-like? I ask because I cannot distinguish all the benchmarks by heart.
François Chollet, creator of ARC-AGI, has consistently said that solving the benchmark does not mean we have AGI. It has always been meant as a stepping stone to encourage progress in the correct direction rather than as an indicator of reaching the destination. That's why he is working on ARC-AGI-3 (to be released in a few weeks) and ARC-AGI-4. His definition of reaching AGI, as I understand it, is when it becomes i…
ARC 2 had a very similar launch.
Both have been crushed in far less time without significantly different architectures than he predicted.
It’s a hard test! And novel, and worth continuing to iterate on. But it was not launched with the humility your last sentence describes.
Re: Gemini 3 Deep Think
#650Earlier quoted context omitted.
Isn’t there? I mean, Claude code has been my biggest usecase and it basically one shots everything now
Yes, LLMs have become extremely good at coding (not software engineer though). But try using them for anything original that cannot be adapted from GitHub and Stack Overflow. I haven't seen much improvement at all at such tasks.
The amount of information available online about optics is probably The gains are likely coming from exactly where they say they are coming from - scaling compute.