Live data from Hacker News

Gemini 2.5 Deep Think

blog.google

211–220 of 259 posts

Re: Gemini 2.5 Deep Think

#211
post #190
post #59

Earlier quoted context omitted.

They might not have been ready/optimized for production, but still wanted to release it before Aug 2 EU AI Act, this way they have 2 years for compliance. So the strategy with aggressively rate-limit for few users make sense.

wheee, great way to lock in incumbents even more or lock out the EU from startups

Welcome to the lovely world of regulation, enjoy your stay.

Re: Gemini 2.5 Deep Think

#212

I started doing some experimentation with this new Deep Think agent, and after five prompts I reached my daily usage limit. For $250 USD/mo that’s what you’ll be getting folks. It’s just bizarrely uncompetitive with o3-pro and Grok 4 Heavy. Anecdotally (from my experience) this was the one feature that enthusiasts in the AI community were interested in to justify the exorbitant price of Google’s Ultra subscription. I…

I'd be interested in tests involving tasks with large amounts of context. Parallel thinking could conceivably useful for a variety of specific problem types. Having more context than any specific chain of thought can reasonably attend to might be one of them.

Re: Gemini 2.5 Deep Think

#214

Earlier quoted context omitted.

Because they want to have their cake and eat it too. A fair pricing model would be token-based, so that a user can see for each query how much they cost, and only pay for what they actually used. But AI companies want a steady stream of income, and they want users to pay as much as possible, while using as little as possible. Therefore they ask for a monthly or even yearly price with an unknown number of tokens inclu…

Personally, I prefer having a fixed, predictable, price rather than paying for usage. There is something psychologically nicer about it to me, and I find myself rationing my usage more when I am using the API (which is effectively what you describe already, just minus the UI).

Yep, this is also why gyms don’t charge you $5 per visit. Nobody would come. Even if it’s cheaper for the average person

Re: Gemini 2.5 Deep Think

#215
post #95
post #78

Earlier quoted context omitted.

"I'm sorry but that wasn't a very interesting question you just asked. I'll spare you the credit and have a cheaper model answer that for you for free. Come back when you have something actually challenging."

Actually why not? Recognizing problem complexity as a fist step is really crucial for such expensive "experts". Humans do the same. And a question to the knowledgeable: does a simple/stupid question cost more in terms of resources then a complex problem? in terms of power consumption.

> ... Recognizing problem complexity as a first step...

Well, I don't think it's easy or even generally possible to recognize a problem complexity. Imagine you ask for a solution for a simple expressed statement like find an n > 2 where z^n = x^n + y^n. The answer you will receive will be based on a trained model with this well known problem but if it's not in the model it could be impossible to measure its complexity.

Re: Gemini 2.5 Deep Think

#216

Earlier quoted context omitted.

Your post misses the fact that 99% of programming is repetitive plumbing and that the overwhelming majority of developers, even ivy league graduates, suck at coding and problem solving. Thus, AI is a great productivity tool if you know how to use it for the overwhelming majority of problems out there. And it's a boost even for those that are not even good at the craft as well. This whole narrative of "okay but it can…

> 99% of programming is repetitive plumbing Even IF that were true (and I'd argue that it is NOT, and it's people who believe that and act that way who produce the tangled messes of spiderweb code that are utterly opaque to public searches and AI analysis -- the supposed "1%"), if even as low as 1% of the code I interacted with was the kind of code that required really deep thought and analysis, it could easily ballo…

Capital is also willing to have vastly lower quality and burden the remaining labor with more toil in exchange for even lower costs. Velocity will rise, quality will fall, toil will increase leading to more burnout but there will be more expendable bodies to cycle through the slop cleanup farm.

Re: Gemini 2.5 Deep Think

#217

Earlier quoted context omitted.

Any reason to think that the wall will be under the human level?

Off the thousands of responses I have read from the top LLMs in the last couple of years: never seen one that was creative. Throwing writing, coding, problem solving, mathematical questions and what not. It's somewhat easier to perceive the creativeless aspect with stable diffusion. I'm not talking about the missing limb or extra finger glitches. With a bit of experience looking through generated images our brain eve…

It doesn't need to be creative, it needs to brute force boilerplate.

Re: Gemini 2.5 Deep Think

#218
post #108

Earlier quoted context omitted.

OK that is recognizably a pelican, pretty great!

This feels like the best pelicanbike yet. The singularity might be closer than we imagine. Time for a leaderboard?

I think they (LLMs providers) are manually tuning these cases/examples.

Pelinkan on a bike - > some dude (from these labs) creates it, and it becomes part of the training data.

Re: Gemini 2.5 Deep Think

#219
post #63
post #57

Been using Gemini for a few months, somehow it's gotten much, much worse in that time. Hallucinations are very common, and it will argue with you when you point it out. So, don't have much confidence.

In my experience with chat, Flash has gotten much, much better. It's my go-to model even though I'm paying for Pro. Pro is frustrating because it too often won't search to find current information, and just gives stale results from before its training cutoff. Flash doesn't do this much anymore. For coding I use Pro in Gemini CLI. It is amazing at coding, but I'm actually using it more to write design docs, decomp mul…

> Flash has gotten much, much better. It's my go-to model even though I'm paying for Pro.

Same I think also Pro got worse...

Re: Gemini 2.5 Deep Think

#220
post #113

I started doing some experimentation with this new Deep Think agent, and after five prompts I reached my daily usage limit. For $250 USD/mo that’s what you’ll be getting folks. It’s just bizarrely uncompetitive with o3-pro and Grok 4 Heavy. Anecdotally (from my experience) this was the one feature that enthusiasts in the AI community were interested in to justify the exorbitant price of Google’s Ultra subscription. I…

Similar complaints are happening all over reddit with the Claude Code $200/mo plan and Cursor. The companies with deep VC funding have been subsidizing usage for a year now, but we're starting to see that bleed off. I think the primary concern of this industry right now is how, relative to the current latest generation models, we simultaneously need intelligence to increase, cost to decrease, effective context window…

> Similar complaints are happening all over reddit with the Claude Code $200/mo

I would imagine 95% of people never get anywhere near to hitting their CC usage. The people who are getting rate-limited have ten windows open, are auto-accepting edits, and YOLO'ing any kind of coherent code quality in their codebase.

Post reply on HN