Live data from Hacker News

Gemini 3 Deep Think

blog.google

121–130 of 722 posts

Re: Gemini 3 Deep Think

#121

Earlier quoted context omitted.

The puzzles are calibrated for human solve rates, but otherwise I agree.

My two elderly parents cannot solve Arc-AGI puzzles, but can manage to navigate the physical world, their house, garden, make meals, clean the house, use the TV, etc. I would say they do have "general intelligence", so whatever Arc-AGI is "solving" it's definitely not "AGI"

You are confusing fluid intelligence with crystallised intelligence.

Re: Gemini 3 Deep Think

#122

Gemini was awesome and now it’s garbage. It’s impossible for it to do anything but cut code down, drop features, lose stuff and give you less than the code you put in. It’s puzzling because it spent months at the head of the pack now I don’t use it at all because why do I want any of those things when I’m doing development. I’m a paid subscriber but there’s no point any more I’ll spend the money on Claude 4.6 instead…

I never found it useful for code. It produced garbage littered with gigantic comments. Me: Remove comments Literally Gemini: // Comments were removed

It would make more sense to me if it had never been awesome.

Re: Gemini 3 Deep Think

#123
post #68

Earlier quoted context omitted.

Yes, but benchmarks like this are often flawed because leading model labs frequently participate in 'benchmarkmaxxing' - ie improvements on ARC-AGI2 don't necessarily indicate similar improvements in other areas (though it does seem like this is a step function increase in intelligence for the Gemini line of models)

Isn’t the point of ARC that you can’t train against it? Or doesn’t it achieve that goal anymore somehow?

How can you make sure of that? AFAIK, these SOTA models run exclusively on their developers hardware. So any test, any benchmark, anything you do, does leak per definition. Considering the nature of us humans and the typical prisoners dilemma, I don't see how they wouldn't focus on improving benchmarks even when it gets a bit... shady?

I tell this as a person who really enjoys AI by the way.

Re: Gemini 3 Deep Think

#126

Earlier quoted context omitted.

The reflection of the sun in the water is completely wrong. LLMs are still useless. (/s)

It's not actually, look up some photos of the sun setting over the ocean. Here's an example: https://stockcake.com/i/sunset-over-ocean_1317824_81961

That’s only if the sun is above the horizon entirely.

Re: Gemini 3 Deep Think

#127

Arc-AGI-2: 84.6% (vs 68.8% for Opus 4.6) Wow. https://blog.google/innovation-and-ai/models-and-research/ge...

Even before this, Gemini 3 has always felt unbelievably 'general' for me. It can beat Balatro (ante 8) with text description of the game alone[0]. Yeah, it's not an extremely difficult goal for humans, but considering:

1. It's an LLM, not something trained to play Balatro specifically

2. Most (probably >99.9%) players can't do that at the first attempt

3. I don't think there are many people who posted their Balatro playthroughs in text form online

I think it's a much stronger signal of its 'generalness' than ARC-AGI. By the way, Deepseek can't play Balatro at all.

[0]: https://balatrobench.com/

Re: Gemini 3 Deep Think

#128
post #15

Google is absolutely running away with it. The greatest trick they ever pulled was letting people think they were behind.

But wait two hours for what OpenAI has! I love the competition and how someone just a few days ago was telling how ARC-AGI-2 was proof that LLMs can't reason. The goalposts will shift again. I feel like most of human endeavor will soon be just about trying to continuously show that AI's don't have AGI.

Soon they can drop the bioweapon to welcome our replacement.

Re: Gemini 3 Deep Think

#129
post #70

I can't shake of the feeling that Googles Deep Think Models are not really different models but just the old ones being run with higher number of parallel subagents, something you can do by yourself with their base model and opencode.

And after i do that, how do i combine the output of 1000 subagents into one output? (Im not being snarky here, i think it's a nontrivial problem)

The idea is that each subagent is focused on a specific part of the problem and can use its entire context window for a more focused subtask than the overall one. So ideally the results arent conflicting, they are complimentary. And you just have a system that merges them.. likely another agent.

Re: Gemini 3 Deep Think

#130
post #15

Google is absolutely running away with it. The greatest trick they ever pulled was letting people think they were behind.

Peacetime Google is not like wartime Google.

Peacetime Google is slow, bumbling, bureaucratic. Wartime Google gets shit done.

Post reply on HN