Live data from Hacker News

Gemini 2.5 Deep Think

blog.google

221–230 of 259 posts

Re: Gemini 2.5 Deep Think

#221

I started doing some experimentation with this new Deep Think agent, and after five prompts I reached my daily usage limit. For $250 USD/mo that’s what you’ll be getting folks. It’s just bizarrely uncompetitive with o3-pro and Grok 4 Heavy. Anecdotally (from my experience) this was the one feature that enthusiasts in the AI community were interested in to justify the exorbitant price of Google’s Ultra subscription. I…

it turns out that AI at this level is very expensive to run (capex, energy). my bet is that AI itself won't figure out how to overcome these constraints and reach escape velocity.

> it turns out that AI at this level is very expensive to run (capex, energy)

If it's CapEx it's -- by definition -- not a cost to run. Energy costs will trend to zero.

Re: Gemini 2.5 Deep Think

#222

I started doing some experimentation with this new Deep Think agent, and after five prompts I reached my daily usage limit. For $250 USD/mo that’s what you’ll be getting folks. It’s just bizarrely uncompetitive with o3-pro and Grok 4 Heavy. Anecdotally (from my experience) this was the one feature that enthusiasts in the AI community were interested in to justify the exorbitant price of Google’s Ultra subscription. I…

Uncompetitive how, what task and Eval?

Gemini is consistently the only model that can reason over long context in dynamic domains for me. Deep Think just did that reviewing an insane amount of Claude Code logs - for a meta analysis task of the underlying implementation. Laughable to think Grok could do that.

Re: Gemini 2.5 Deep Think

#223
post #63
post #57

Been using Gemini for a few months, somehow it's gotten much, much worse in that time. Hallucinations are very common, and it will argue with you when you point it out. So, don't have much confidence.

In my experience with chat, Flash has gotten much, much better. It's my go-to model even though I'm paying for Pro. Pro is frustrating because it too often won't search to find current information, and just gives stale results from before its training cutoff. Flash doesn't do this much anymore. For coding I use Pro in Gemini CLI. It is amazing at coding, but I'm actually using it more to write design docs, decomp mul…

interesting out of all "thinking models," I struggle with Gemini the most for coding. Just can't make it perform. I feel like they silently nerfed it over the last months.

Re: Gemini 2.5 Deep Think

#224
post #95
post #78

Earlier quoted context omitted.

"I'm sorry but that wasn't a very interesting question you just asked. I'll spare you the credit and have a cheaper model answer that for you for free. Come back when you have something actually challenging."

Actually why not? Recognizing problem complexity as a fist step is really crucial for such expensive "experts". Humans do the same. And a question to the knowledgeable: does a simple/stupid question cost more in terms of resources then a complex problem? in terms of power consumption.

Google actually does provide that service! https://cloud.google.com/vertex-ai/generative-ai/docs/model-...

Re: Gemini 2.5 Deep Think

#225

Earlier quoted context omitted.

Several years ago I thought a good litmus test for mastery of coding is not finding a solution using internet search nor getting well written questions about esoteric coding problems answered on StackOverflow. For a while, I would post a question and answer my own question after I solved the problem for posterity (or AI bots). I always loved getting the "I've been working on this for 3 days and you saved my life" com…

They're remarkably useless on stuff they've seen but not had up-weighted in the training set. Even the best ones (Opus 4 running hot, Qwen and K2 will surprise you fairly often) are a net liability in some obscure thing. Probably the starkest example of this is build system stuff: it's really obvious which ones have seen a bunch of `nixpkgs`, and even the best ones seem to really struggle with Bazel and sometimes CMa…

Yeah, I mean the interpolation part is new, but boy do I miss pre-enshitification Google!

Re: Gemini 2.5 Deep Think

#226

Earlier quoted context omitted.

Several years ago I thought a good litmus test for mastery of coding is not finding a solution using internet search nor getting well written questions about esoteric coding problems answered on StackOverflow. For a while, I would post a question and answer my own question after I solved the problem for posterity (or AI bots). I always loved getting the "I've been working on this for 3 days and you saved my life" com…

Your post misses the fact that 99% of programming is repetitive plumbing and that the overwhelming majority of developers, even ivy league graduates, suck at coding and problem solving. Thus, AI is a great productivity tool if you know how to use it for the overwhelming majority of problems out there. And it's a boost even for those that are not even good at the craft as well. This whole narrative of "okay but it can…

I've started to come to the conclusion that only greenfield projects consist of repetitive plumbing. Legacy software is like plumbing if all the pipes were tied into a knot. The edge cases, ambiguous naming, hacky solutions, etc. all make for a miserable experience, both for humans and AIs.

Re: Gemini 2.5 Deep Think

#227

Earlier quoted context omitted.

it turns out that AI at this level is very expensive to run (capex, energy). my bet is that AI itself won't figure out how to overcome these constraints and reach escape velocity.

> it turns out that AI at this level is very expensive to run (capex, energy) If it's CapEx it's -- by definition -- not a cost to run. Energy costs will trend to zero.

Why will energy costs trend to zero?

Re: Gemini 2.5 Deep Think

#228

Earlier quoted context omitted.

> it turns out that AI at this level is very expensive to run (capex, energy) If it's CapEx it's -- by definition -- not a cost to run. Energy costs will trend to zero.

Why will energy costs trend to zero?

Renewables will keep getting more efficient and cheaper to install, batteries will continue to get cheaper, at some point they'll crack fusion. Prices go negative or zero in several places already (West Texas wind energy overnight, solar in Chile). The question seems less "when will we get abundant, virtually free clean energy" and more "will we do it in time to avoid climate collapse".

Re: Gemini 2.5 Deep Think

#229

Earlier quoted context omitted.

Any reason to think that the wall will be under the human level?

Off the thousands of responses I have read from the top LLMs in the last couple of years: never seen one that was creative. Throwing writing, coding, problem solving, mathematical questions and what not. It's somewhat easier to perceive the creativeless aspect with stable diffusion. I'm not talking about the missing limb or extra finger glitches. With a bit of experience looking through generated images our brain eve…

An opinion on the current state of the field. The usual stochastic parrot mention. That, I see. Reasons for the existence of the wall? Not so much.

Re: Gemini 2.5 Deep Think

#230
post #152

Earlier quoted context omitted.

This could fix my main gripe with The Matrix. ”Humans are used as batteries” always felt off, but it totally would make sense if the human brains have uniquely energy efficient pattern matching abilities that an emerging AI organism would harvest. That would also strengthen the spiritual humanist subtext.

Thats because the original script did actually have the human farms being used for brainpower for the machines. They changed it to "batteries" because they thought audiences at the time wouldnt understand it!

Wow, TIL. That strengthens my perception of the masterpiece. Thanks
Post reply on HN