Live data from Hacker News

Gemini 3 Deep Think

blog.google

621–630 of 722 posts

Re: Gemini 3 Deep Think

#621
post #522

Earlier quoted context omitted.

Then there is the third axis, intelligence. To continue your chain: Eurasian magpies are conscious, but also know themselves in the mirror (the "mirror self-recognition" test). But yet, something is still missing.

The mirror test doesn’t measure intelligence so much as it measures mirror aptitude. It’s prone to over fitting.

Exactly, it's a poor test. Consider the implication that the blind cant be fully conscious.

It's a test of perceptual ability, not introspection.

Re: Gemini 3 Deep Think

#622

Arc-AGI-2: 84.6% (vs 68.8% for Opus 4.6) Wow. https://blog.google/innovation-and-ai/models-and-research/ge...

Am I the only one that can’t find Gemini useful except if you want something cheap? I don’t get what was the whole code red about or all that PR. To me I see no reason to use Gemini instead of of GPT and Anthropic combo. I should add that I’ve tried it as chat bot, coding through copilot and also as part of a multi model prompt generation. Gemini was always the worst by a big margin. I see some people saying it is sm…

Yeah it's pretty shit compared to Opus

Re: Gemini 3 Deep Think

#623

Earlier quoted context omitted.

5 days for Ai is by no mean short! If it can solve it, it would need perhaps 1-2 hours. If it can not, 5 days continuous running would produce gibberish only. We can safely assume that such private models will run inferences entirely on dedicated hardware, sharing with nobody. So if they could not solve the problems, it's not due to any artificial constraint or lack of resources, far from it. The 5 days window, howev…

5 days is short for memetic propagation on social media to reach everyone who has their own harness and agentic setup that wants to have a go.

That's not really how it works, the recent Erdos proofs in Lean were done by a specialized proprietary model (Aristotle by Harmonic) that's specifically trained for this task. Normal agents are not effective.

Re: Gemini 3 Deep Think

#624
post #425

Earlier quoted context omitted.

Israel is not one of the boots. Deplorable as their domestic policy may be, they're not wagging the dog of capitalist imperialism. To imply otherwise is to reveal yourself as biased, warped in a way that keeps you from going after much bigger, and more real systems of political economy holding back our civilization from universal human dignity and opportunity.

Lol what? Not sure if you are defending Israel or google because your communication style is awful. But if you are defending Israel then you're an idiot who is excusing genocide. If you're defending google then you're just a corporate bootlicker who means nothing.

You edited your comment.

Re: Gemini 3 Deep Think

#625

Earlier quoted context omitted.

My experience with Antigravity is the opposite. It's the first time in over 10 years that an IDE has managed to take me out a bit out of the jetbrain suite. I did not think that was something possible as I am a hardcore jetbrain user/lover.

It's literally just vscode? I tried it the other day and I couldn't tell it apart from windsurf besides the icon in my dock

Yeah same here. Even though it's vscode I'm still using it and don't plan to renew Intellij again. Gemini was crap but Opus smashes it.

It is windsurf isn't it, why would you expect it to be different?

Re: Gemini 3 Deep Think

#626

Earlier quoted context omitted.

Agreed, it's a truly wild take. While I fully support the humility of not knowing, at a minimum I think we can say determinations of consciousness have some relation to specific structure and function that drive the outputs, and the actual process of deliberating on whether there's consciousness would be a discussion that's very deep in the weeds about architecture and processes. What's fascinating is that evolution…

> at a minimum I think we can say determinations of consciousness have some relation to specific structure and function that drive the outputs Every time anyone has tried that it excludes one or more classes of human life, and sometimes led to atrocities. Let's just skip it this time.

I excluded all right handed, blue eyed people yesterday before breakfast. No atrocities happened because of it.

Re: Gemini 3 Deep Think

#627

Earlier quoted context omitted.

> their Deep Research usually works the hardest That's sortof damning with faint praise I think. So, for $work I needed to understand the legal landscape for some regulations (around employment screening) so I kicked off a deep research for all the different countries. That was fineish, but tended to go off the rails towards the end. So, then I split it out into Americas, APAC and EMEA requirements. This time, I spen…

Oh yeah, LLMs currently spew a lot of garbage. Everything has to be double-checked. I mainly use them for gathering sources and pointing out a few considerations I might have otherwise overlooked. I often run them a few times, because they go off the rails in different directions, but sometimes those directions are helpful for me in expanding my understanding. I still have to synthesize everything from scratch myself…

> For me it's less about saving time, and more about potentially unearthing good sources that my google searches wouldn't turn up, and occasionally giving me a few nuggets of inspiration / new rabbit holes to go down.

Yeah, I see the value here. And for personal stuff, that's totally fine. But these tools are being sold to businesses as productivity increasers, and I'm not buying it right now.

I really, really want this to work though, as it would be such a massive boost to human flourishing. Maybe LLMs are the wrong approach though, certainly the current models aren't doing a good job.

Re: Gemini 3 Deep Think

#628
post #603

Earlier quoted context omitted.

It depends on time. 5 years ago it was quite well defined that it’s the last one, maybe the second one in some context. Especially when distinction was important, it was always the last one. In our case it was. We trained models to have weights. We even stored models and weights separately, because models change slower than weights. You could choose a model and a set of weights, and run them. You could change weights…

It seems unlikely "model" was ever equivalent in meaning to "architecture". Otherwise there would be just one "CNN model" or just one "transformer model" insofar there is a single architecture involved.

First of all, hyperparameters. Second, organization, or connections. 3rd, cost function. 4th, activation function. 5th type of learning. Etc.

These are not weights. These were parts of models.

Re: Gemini 3 Deep Think

#629
post #626

Earlier quoted context omitted.

> at a minimum I think we can say determinations of consciousness have some relation to specific structure and function that drive the outputs Every time anyone has tried that it excludes one or more classes of human life, and sometimes led to atrocities. Let's just skip it this time.

I excluded all right handed, blue eyed people yesterday before breakfast. No atrocities happened because of it.

And people say the machines don't learn!

Re: Gemini 3 Deep Think

#630
post #15

Google is absolutely running away with it. The greatest trick they ever pulled was letting people think they were behind.

They seem to be optimizing for benchmarks instead of real world use

Yeah if only Gemini performed half as well as it does on benches, we'd actually be using it.
Post reply on HN