Live data from Hacker News

Gemini 3 Deep Think

blog.google

561–570 of 722 posts

Re: Gemini 3 Deep Think

#561

Earlier quoted context omitted.

> If something claims to be conscious and we can't prove it's not, we have no choice but to believe it. This is not a good test. A dog won't claim to be conscious but clearly is, despite you not being able to prove one way or the other. GPT-3 will claim to be conscious and (probably) isn't, despite you not being able to prove one way or the other.

An LLM will claim whatever you tell it to claim. (In fact this Hacker News comment is also conscious.) A dog won’t even claim to be a good boy.

A classic relevant comic:

https://www.threepanelsoul.com/comic/dog-philosophy

Re: Gemini 3 Deep Think

#562

Earlier quoted context omitted.

Who said they’re godlike today? And yes, you are probably using them wrong if you don’t find them useful or don’t see the rapid improvement.

Let's come back in 12 months and discuss your singularity then. Meanwhile I spent like $30 on a few models as a test yesterday, none of them could tell me why my goroutine system was failing, even though it was painfully obvious (I purposefully added one too many wg.Done), gemini, codex, minimax 2.5, they all shat the bed on a very obvious problem but I am to believe they're 98% conscious and better at logic and math…

It's basically bunch of people who see themselves as too smart to believe in God, instead they have just replaced it with AI and Singularity and attribute similar stuff to it eg. eternal life which is just heaven in religion. Amodei was hawking doubling of human lifespan to a bunch of boomers not too long ago. Ponce de León also went to search for the fountain of youth. It's a very common theme across human history. AI is just the new iteration where they mirror all their wishes and hopes.

Re: Gemini 3 Deep Think

#563
post #280

Earlier quoted context omitted.

What's the point of denying or downplaying that we are seeing amazing and accelerating advancements in areas that many of us thought were impossible?

It can be reasonable to be skeptical that advances on benchmarks may be only weakly or even negatively correlated with advances on real-world tasks. I.e. a huge jump on benchmarks might not be perceptible to 99% of users doing 99% of tasks, or some users might even note degradation on specific tasks. This is especially the case when there is some reason to believe most benchmarks are being gamed. Real-world use is wh…

The GP comment is not skeptical of the jump in benchmark scores reported by one particular LLM. It's skeptical of machine intelligence in general, claims that there's no value in comparing their performances with those of human beings, and accuses those who disagree with this take of "hubris and grift". This has nothing to do with any form or reasonable skepticism.

Re: Gemini 3 Deep Think

#564
post #507

Earlier quoted context omitted.

Disagree. Claude Code is great for coding, Gemini is better than everything else for everything else.

What is "everything else" in your view? Just curious -- I really only seriously use models for coding, so I am curious what I am missing.

Role-playing but Claude is as bad, same censored garbage with the CEO wanting to be your dad. Grok is best for everything else by far.

Re: Gemini 3 Deep Think

#565
post #382

So what happens if the AI companies can't make money? I see more and more advances and breakthrough but they are taking in debt and no revenue in sight. I seem to understand debt is very bad here since they could just sell more shares, but aren't (either valuation is stretched or no buyers). Just a recession? Something else? Aren't they very very big to fall? Edit0: Revenue isn't the right word, profit is more correc…

What happens if oil companies can't make money? They will restructure society so they can. That's the essence of capitalism, the willingness to restructure society to chase growth.

Obviously this tech is profitable in some world. Car companies can't make money if we live in walking distance and people walk on roads.

Re: Gemini 3 Deep Think

#566

Earlier quoted context omitted.

Don't let the benchmarks fool you. Gemini models are completely useless not matter how smart they are. Google still hasn't figure out tool calling and making the model follow instructions. They seem to only care about benchmarking and being the most intelligent model on paper. This has been a problem of Gemini since 1.0 and they still haven't fixed it. Also the worst model in terms of hallucinations.

Disagree. Claude Code is great for coding, Gemini is better than everything else for everything else.

And mathematics?

Re: Gemini 3 Deep Think

#567
post #53

Earlier quoted context omitted.

https://arcprize.org/leaderboard $13.62 per task - so we need another 5-10 years for the price to run this to become reasonable? But the real question is if they just fit the model to the benchmark.

A grad student hour is probably more expensive…

As it should be. They're a human!

Re: Gemini 3 Deep Think

#568

Earlier quoted context omitted.

> His definition of reaching AGI, as I understand it, is when it becomes impossible to construct the next version of ARC-AGI because we can no longer find tasks that are feasible for normal humans but unsolved by AI. That is the best definition I've yet to read. If something claims to be conscious and we can't prove it's not, we have no choice but to believe it. Thats said, I'm reminded of the impossible voting tests…

>because we can no longer find tasks that are feasible for normal humans but unsolved by AI. "Answer "I don't know" if you don't know an answer to one of the questions"

Normal humans don't pass this benchmark either, as evidenced by the existence of religion, among other things.

Re: Gemini 3 Deep Think

#569

Earlier quoted context omitted.

The 1st proof original solutions are due to be published in about 24h, AIUI.

Feels like an unforced blunder to make the time window so short after going to so much effort and coming up with something so useful.

5 days for Ai is by no mean short! If it can solve it, it would need perhaps 1-2 hours. If it can not, 5 days continuous running would produce gibberish only. We can safely assume that such private models will run inferences entirely on dedicated hardware, sharing with nobody. So if they could not solve the problems, it's not due to any artificial constraint or lack of resources, far from it.

The 5 days window, however, is a sweat spot because it likely prevents cheating by hiring a math PhD and feed the AI with hints and ideas.

Re: Gemini 3 Deep Think

#570
post #291

Earlier quoted context omitted.

Burned in seconds.

Getting the work done faster for the same money doesn't make the work more expensive. You could slow down the inference to make the task take longer, if $/sec matters.

You're right, but I don't think we're getting an hour's worth of work out of single prompts yet. Usually it's an hour's worth of work out of 10 prompts for iteration. Now that's a day's wage for an hour of work. I'm certain the crossover will come soon, but it doesn't feel there yet.
Post reply on HN