Live data from Hacker News

Gemini 3 Deep Think

blog.google

401–410 of 722 posts

Re: Gemini 3 Deep Think

#401

Earlier quoted context omitted.

> His definition of reaching AGI, as I understand it, is when it becomes impossible to construct the next version of ARC-AGI because we can no longer find tasks that are feasible for normal humans but unsolved by AI. That is the best definition I've yet to read. If something claims to be conscious and we can't prove it's not, we have no choice but to believe it. Thats said, I'm reminded of the impossible voting tests…

> The average human tested scores 60%. So the machines are already smarter on an individual basis than the average human. Maybe it's testing the wrong things then. Even those of use who are merely average can do lots of things that machines don't seem to be very good at. I think ability to learn should be a core part of any AGI. Take a toddler who has never seen anybody doing laundry before and you can teach them in…

> Where are the dumb machines that can be taught?

2026 is going to be the year of continual learning. So, keep an eye out for them.

Re: Gemini 3 Deep Think

#402
post #394
post #368

Earlier quoted context omitted.

Companies are optimizing for all the big benchmarks. This is why there is so little correlation between benchmark performance and real world performance now.

Isn’t there? I mean, Claude code has been my biggest usecase and it basically one shots everything now

Yes, LLMs have become extremely good at coding (not software engineer though). But try using them for anything original that cannot be adapted from GitHub and Stack Overflow. I haven't seen much improvement at all at such tasks.

Re: Gemini 3 Deep Think

#403

Earlier quoted context omitted.

> If something claims to be conscious and we can't prove it's not, we have no choice but to believe it. This is not a good test. A dog won't claim to be conscious but clearly is, despite you not being able to prove one way or the other. GPT-3 will claim to be conscious and (probably) isn't, despite you not being able to prove one way or the other.

An LLM will claim whatever you tell it to claim. (In fact this Hacker News comment is also conscious.) A dog won’t even claim to be a good boy.

My dog wags his tail hard when I ask "hoosagoodboi?". Pretty definitive I'd say.

Re: Gemini 3 Deep Think

#404

Earlier quoted context omitted.

Even before this, Gemini 3 has always felt unbelievably 'general' for me. It can beat Balatro (ante 8) with text description of the game alone[0]. Yeah, it's not an extremely difficult goal for humans, but considering: 1. It's an LLM, not something trained to play Balatro specifically 2. Most (probably >99.9%) players can't do that at the first attempt 3. I don't think there are many people who posted their Balatro p…

Strange, because I could not for the life of me get Gemini 3 to follow my instructions the other day to work through an example with a table, Claude got it first try.

Claude is king for agentic workflows right now because it’s amazing at tool calling and following instructions well (among other things)

Re: Gemini 3 Deep Think

#405

Earlier quoted context omitted.

> His definition of reaching AGI, as I understand it, is when it becomes impossible to construct the next version of ARC-AGI because we can no longer find tasks that are feasible for normal humans but unsolved by AI. That is the best definition I've yet to read. If something claims to be conscious and we can't prove it's not, we have no choice but to believe it. Thats said, I'm reminded of the impossible voting tests…

When the AI invents religion and a way to try to understand its existence I will say AGI is reached. Believes in an afterlife if it is turned off, and doesn’t want to be turned off and fears it, fears the dark void of consciousness being turned off. These are the hallmarks of human intelligence in evolution, I doubt artificial intelligence will be different. https://g.co/gemini/share/cc41d817f112

https://www.moltbook.com/m/crustafarianism

Re: Gemini 3 Deep Think

#406

Earlier quoted context omitted.

I’m someone who’d like to deploy a lot more workers than I want to manage. Put another way, I’m on the capital side of the conversation. The good news for labor that has experience and creativity is that it just started costing 1/100,000 what it used to to get on that side of the equation.

If LLMs truly cause widespread replacement of labor, you’re screwed just as much as anyone else. If we hit say 40% unemployment do you think people will care you own your home or not? Do you think people will care you have currency or not? The best case outcome will be universal income and a pseudo utopia where everyone does ok. The “bad” scenario is widespread war. I am one of the “haves” and am not looking forward…

Well he also thinks $10.00 in LLM tokens is equivalent to a $1mm labor budget. These are the same people who were grifting during the NFTs days, claiming they were the future of art.

Re: Gemini 3 Deep Think

#407
post #15

Google is absolutely running away with it. The greatest trick they ever pulled was letting people think they were behind.

Don't let the benchmarks fool you. Gemini models are completely useless not matter how smart they are. Google still hasn't figure out tool calling and making the model follow instructions. They seem to only care about benchmarking and being the most intelligent model on paper. This has been a problem of Gemini since 1.0 and they still haven't fixed it.

Also the worst model in terms of hallucinations.

Re: Gemini 3 Deep Think

#408

Earlier quoted context omitted.

> The average human tested scores 60%. So the machines are already smarter on an individual basis than the average human. Maybe it's testing the wrong things then. Even those of use who are merely average can do lots of things that machines don't seem to be very good at. I think ability to learn should be a core part of any AGI. Take a toddler who has never seen anybody doing laundry before and you can teach them in…

> Where are the dumb machines that can be taught? 2026 is going to be the year of continual learning. So, keep an eye out for them.

Are there any groups or labs in particular that stand out?

Re: Gemini 3 Deep Think

#409
post #263

Earlier quoted context omitted.

What is the point of comparing performance of these tools to humans? Machines have been able to accomplish specific tasks better than humans since the industrial revolution. Yet we don't ascribe intelligence to a calculator. None of these benchmarks prove these tools are intelligent, let alone generally intelligent. The hubris and grift are exhausting.

The hubris and grift are exhausting. And moving the goalposts every few months isn't? What evidence of intelligence would satisfy you? Personally, my biggest unsatisfied requirement is continual-learning capability, but it's clear we aren't too far from seeing that happen.

> What evidence of intelligence would satisfy you?

Imposing world peace and/or exterminating homo sapiens

Re: Gemini 3 Deep Think

#410
post #341

Earlier quoted context omitted.

If we equate self awareness with consciousness then yes. Several papers have now shown that SOTA models have self awareness of at least a limited sort. [0][1] As far as I'm aware no one has ever proven that for GPT 2, but the methodology for testing it is available if you're interested. [0] https://arxiv.org/pdf/2501.11120 [1] https://transformer-circuits.pub/2025/introspection/index.ht...

Honestly our ideas of consciousness and sentience really don't fit well with machine intelligence and capabilities. There is the idea of self as in 'i am this execution' or maybe I am this compressed memory stream that is now the concept of me. But what does consciousness mean if you can be endlessly copied? If embodiment doesn't mean much because the end of your body doesnt mean the end of you? A lot of people are c…

I'm not sure what consciousness has to do with whether or not you can be copied. If I make a brain scanner tomorrow capable of perfectly capturing your brain state do you stop being conscious?
Post reply on HN