Live data from Hacker News

Gemini 3 Deep Think

blog.google

641–650 of 722 posts

Re: Gemini 3 Deep Think

#641
post #148

Earlier quoted context omitted.

I look forward to them trying. I'll know when the pelican riding a bicycle is good but the ocelot riding a skateboard sucks.

Would it not be better to have 100 such tests "Pelican on bicycle", "Tiger on stilts"..., and generate them all for every new model but only release a new one each time. That way you could show progression across all models, attempts at benchmaxxing would be more obvious. Given the crazy money and vying for supremacy among AI companies right now it does seem naive to belive that no attempt at better pelicans on bicyc…

Or indeed do the Markov chain conceptual slip. Pelican on bicycle, badger on stool, tiger on acid. Pelican on bicycle is definitely cooked, though: people know it and it's talked about in language.

Re: Gemini 3 Deep Think

#642
post #39

The pelican riding a bicycle is excellent . I think it's the best I've seen. https://simonwillison.net/2026/Feb/12/gemini-3-deep-think/

Do you have to still keep trying to bang on about this relentlessly? It was sort of humorous for the maybe first 2 iterations, now it's tacky, cheesy, and just relentless self-promotion. Again, like I said before, it's also a terrible benchmark.

It's HN's Carthago delenda est moment.

Re: Gemini 3 Deep Think

#643

Earlier quoted context omitted.

Wartime Google gave us Google+. Wartime Google is still bumbling, and despite OpenAI's numerous missteps, I don't think it has to worry about Google hurting its business yet.

Google+ was fun. Failed in the market though. Apple made a social network called Ping. Disaster. MobileMe was silly. Microsoft made Zune and the Kin 1 and Kin 2 devices and Windows phone and all sorts of other disasters. These things happen.

I have a hypothesis that Google+ just wasn't addictive. Which is a good thing now, but not back then

Re: Gemini 3 Deep Think

#644

Earlier quoted context omitted.

That's not really how it works, the recent Erdos proofs in Lean were done by a specialized proprietary model (Aristotle by Harmonic) that's specifically trained for this task. Normal agents are not effective.

Why did you omit the other AI-generated Erdos proofs not done by a proprietary model, which occurred on timescales stretched across significantly longer time than 5 days?

Those were not really "proofs" by the standard of 1stproof. The only way an AI can possibly convince an unsympathetic peer reviewer that its proof is correct is to write it completely in a formal system like Lean. The so-called "proofs" done with GPT were half baked and required significant human input, hints, fixing after the fact etc. which is enough to disqualify them from this effort.

Re: Gemini 3 Deep Think

#645

Earlier quoted context omitted.

> If something claims to be conscious and we can't prove it's not, we have no choice but to believe it. This is not a good test. A dog won't claim to be conscious but clearly is, despite you not being able to prove one way or the other. GPT-3 will claim to be conscious and (probably) isn't, despite you not being able to prove one way or the other.

An LLM will claim whatever you tell it to claim. (In fact this Hacker News comment is also conscious.) A dog won’t even claim to be a good boy.

This isn't really as true anymore.

Last week gemini argued with me about an auxiliary electrical generator install method and it turned out to be right, even though I pushed back hard on it being incorrect. First time that has ever happened.

Re: Gemini 3 Deep Think

#646
post #601
post #346

Earlier quoted context omitted.

>Why is it so easy for me to open the car door Because this part of your brain has been optimized for hundreds of millions of years. It's been around a long ass time and takes an amazingly low amount of energy to do these things. On the other hand the 'thinking' part of your brain, that is your higher intelligence is very new to evolution. It's expensive to run. It's problematic when giving birth. It's really slow wi…

> There's a term for this, but I can't think of it at the moment. Moravec's paradox: https://epoch.ai/gradient-updates/moravec-s-paradox

Thanks, I can never quite remember that.

Re: Gemini 3 Deep Think

#647

Earlier quoted context omitted.

François Chollet, creator of ARC-AGI, has consistently said that solving the benchmark does not mean we have AGI. It has always been meant as a stepping stone to encourage progress in the correct direction rather than as an indicator of reaching the destination. That's why he is working on ARC-AGI-3 (to be released in a few weeks) and ARC-AGI-4. His definition of reaching AGI, as I understand it, is when it becomes i…

> His definition of reaching AGI, as I understand it, is when it becomes impossible to construct the next version of ARC-AGI because we can no longer find tasks that are feasible for normal humans but unsolved by AI. That is the best definition I've yet to read. If something claims to be conscious and we can't prove it's not, we have no choice but to believe it. Thats said, I'm reminded of the impossible voting tests…

Wait where does the idea of consciousness enter this? AGI doesn't need to be conscious.

Re: Gemini 3 Deep Think

#648
post #183

Earlier quoted context omitted.

The idea that an AI lab would pay a small army of human artists to create training data for $animal on $transport just to cheat on my stupid benchmark delights me.

When you're spending trillions on capex, paying a couple of people to make some doodles in SVGs would not be a big expense.

I think no matter what happens with AI in the future, there will always be a subset of people with elaborate conspiracies about how it's all fake/a hoax.

Re: Gemini 3 Deep Think

#649
post #26

Earlier quoted context omitted.

Weren't we barely scraping 1-10% on this with state of the art models a year ago and it was considered that this is the final boss, ie solve this and its almost AGI-like? I ask because I cannot distinguish all the benchmarks by heart.

François Chollet, creator of ARC-AGI, has consistently said that solving the benchmark does not mean we have AGI. It has always been meant as a stepping stone to encourage progress in the correct direction rather than as an indicator of reaching the destination. That's why he is working on ARC-AGI-3 (to be released in a few weeks) and ARC-AGI-4. His definition of reaching AGI, as I understand it, is when it becomes i…

Please let’s hold M Chollet to account, at least a little. He launched ARC claiming transformer architectures could never do it and that he thought solving it would be AGI. And he was smug about it.

ARC 2 had a very similar launch.

Both have been crushed in far less time without significantly different architectures than he predicted.

It’s a hard test! And novel, and worth continuing to iterate on. But it was not launched with the humility your last sentence describes.

Re: Gemini 3 Deep Think

#650
post #402
post #394

Earlier quoted context omitted.

Isn’t there? I mean, Claude code has been my biggest usecase and it basically one shots everything now

Yes, LLMs have become extremely good at coding (not software engineer though). But try using them for anything original that cannot be adapted from GitHub and Stack Overflow. I haven't seen much improvement at all at such tasks.

No shot, their classic engineering ability has exploded too.

The amount of information available online about optics is probably The gains are likely coming from exactly where they say they are coming from - scaling compute.

Post reply on HN