Live data from Hacker News

Gemini 3 Deep Think

blog.google

251–260 of 722 posts

Re: Gemini 3 Deep Think

#251

Less than a year to destroy Arc-AGI-2 - wow.

It's a useless meaningless benchmark though, it just got a catchy name, as in, if the models solve this it means they have "AGI", which is clearly rubbish. Arc-AGI score isn't correlated with anything useful.

It's correlated with the ability to solve logic puzzles.

It's also interesting because it's very very hard for base LLMs, even if you try to "cheat" by training on millions of ARC-like problems. Reasoning LLMs show genuine improvement on this type of problem.

Re: Gemini 3 Deep Think

#252
post #15

Google is absolutely running away with it. The greatest trick they ever pulled was letting people think they were behind.

Gemini's UX (and of course privacy cred as with anything Google) is the worst of all the AI apps. In the eyes of the Common Man, it's UI that will win out, and ChatGPT's is still the best.

Gemini is completely unusable in VS Code. It's rated 2/5 stars, pathetic: https://marketplace.visualstudio.com/items?itemName=Google.g...

Requests regularly time out, the whole window freezes, it gets stuck in schizophrenic loops, edits cannot be reverted and more.

It doesn't even come close to Claude or ChatGPT.

Re: Gemini 3 Deep Think

#253

Earlier quoted context omitted.

Who said they’re godlike today? And yes, you are probably using them wrong if you don’t find them useful or don’t see the rapid improvement.

Let's come back in 12 months and discuss your singularity then. Meanwhile I spent like $30 on a few models as a test yesterday, none of them could tell me why my goroutine system was failing, even though it was painfully obvious (I purposefully added one too many wg.Done), gemini, codex, minimax 2.5, they all shat the bed on a very obvious problem but I am to believe they're 98% conscious and better at logic and math…

Post the file here

Re: Gemini 3 Deep Think

#254
post #240

Earlier quoted context omitted.

I'd rather say it has a mind of its own; it does things its way. But I have not tested this model, so they might have improved its instruction following.

Well, one thing i know for sure: it reliably misplaces parentheses in lisps.

Clearly, the AI is trying to steer you towards the ML family of languages for its better type system, performance, and concurrency ;)

Re: Gemini 3 Deep Think

#256
post #216

Earlier quoted context omitted.

> His definition of reaching AGI, as I understand it, is when it becomes impossible to construct the next version of ARC-AGI because we can no longer find tasks that are feasible for normal humans but unsolved by AI. That is the best definition I've yet to read. If something claims to be conscious and we can't prove it's not, we have no choice but to believe it. Thats said, I'm reminded of the impossible voting tests…

> If something claims to be conscious and we can't prove it's not, we have no choice but to believe it. Can you "prove" that GPT2 isn't concious?

If we equate self awareness with consciousness then yes. Several papers have now shown that SOTA models have self awareness of at least a limited sort. [0][1]

As far as I'm aware no one has ever proven that for GPT 2, but the methodology for testing it is available if you're interested.

[0]https://arxiv.org/pdf/2501.11120

[1]https://transformer-circuits.pub/2025/introspection/index.ht...

Re: Gemini 3 Deep Think

#257

Gemini has always felt like someone who was book smart to me. It knows a lot of things. But if you ask it do anything that is offscript it completely falls apart

I strongly suspect there's a major component of this type of experience being that people develop a way of talking to a particular LLM that's very efficient and works well for them with it, but is in many respects non-transferable to rival models. For instance, in my experience, OpenAI models are remarkably worse than Google models in basically any criterion I could imagine; however, I've spent most of my time using the Google ones and it's only during this time that the differences became apparent and, over time, much more pronounced. I would not be surprised at all to learn that people who chose to primarily use Anthropic or OpenAI models during that time had an exactly analogous experience that convinced them their model was the best.

Re: Gemini 3 Deep Think

#258
post #95
post #74

Earlier quoted context omitted.

Its really weird how you all are begging to be replaced by llms, you think if agentic workflows get good enough you're going to keep your job? Or not have your salary reduced by 50%? If Agents get good enough it's not going to build some profitable startup for you (or whatever people think they're doing with the llm slot machines) because that implies that anyone else with access to that agent can just copy you, its…

[flagged]

[flagged]

Re: Gemini 3 Deep Think

#259

Earlier quoted context omitted.

Could it also be that the models are just a lot better than a year ago?

> Could it also be that the models are just a lot better than a year ago? No, the proof is in the pudding. After AI we're having higher prices, higher deficits and lower standard of living. Electricity, computers and everything else costs more. "Doing better" can only be justified by that real benchmark. If Gemini 3 DT was better we would have falling prices of electricity and everything else at least until they get…

You might call me crazy, but at least in 2024, consumers spent ~1% less of their income on expenses than 2019[2], which suggests that 2024 is more affordable than 2019.

This is from the BLS consumer survey report released in dec[1]

[1]https://www.bls.gov/news.release/cesan.nr0.htm

[2]https://www.bls.gov/opub/reports/consumer-expenditures/2019/

Prices are never going back to 2019 numbers though

Re: Gemini 3 Deep Think

#260
post #151

Earlier quoted context omitted.

Google privacy cred is ... excellent? The worst data breach I know of them having was a flaw that allowed access to names and emails of 500k users.

Link? Are you conflating with "500k Gmail accounts leaked [by a third party]" with Gmail having a breach? Afaik, Google has had no breaches ever.

https://en.wikipedia.org/wiki/2018_Google_data_breach
Post reply on HN