Live data from Hacker News

Gemini 2.5 Deep Think

blog.google

81–90 of 259 posts

Re: Gemini 2.5 Deep Think

#81

I started doing some experimentation with this new Deep Think agent, and after five prompts I reached my daily usage limit. For $250 USD/mo that’s what you’ll be getting folks. It’s just bizarrely uncompetitive with o3-pro and Grok 4 Heavy. Anecdotally (from my experience) this was the one feature that enthusiasts in the AI community were interested in to justify the exorbitant price of Google’s Ultra subscription. I…

Several years ago I thought a good litmus test for mastery of coding is not finding a solution using internet search nor getting well written questions about esoteric coding problems answered on StackOverflow. For a while, I would post a question and answer my own question after I solved the problem for posterity (or AI bots). I always loved getting the "I've been working on this for 3 days and you saved my life" comments.

I've been working on a challenging problem all this week and all the AI copilot models are worthless helping me. Mastery in coding is being alone when nobody else nor AI copilots can help you and you have dig deep into generalization, synthesis, and creativity.

(I thought to myself, at least it will be a little while longer before I'm replaced with AI coding agents.)

Re: Gemini 2.5 Deep Think

#82
post #60

Earlier quoted context omitted.

It doesn't, it's not "1000 Gemini Pro" requests for free, Google misled everyone. It's 1000 Gemini requests, Flash included. You get like 5-7 Gemini Pro requests before you get limited.

I'm getting 100 Gemini Pro requests per day with an AI Studio API key that doesn't have billing enabled. After that it's bumped down to Flash, which is surpisingly effective in Gemini CLI. If I need Pro, I just swap in an API from an account with billing enabled, but usually 100 requests is enough for a day of work.

Whoa. I’m definitely getting just handful of requests via free, just like patent commenter…

Re: Gemini 2.5 Deep Think

#83

Approach is analogous to Grok 4 Heavy: use multiple "reasoning" agents in parallel and then compare answers before coming back with a single response, taking ~30 minutes. Great results, though it would be more fair for the benchmark comparisons to be against Grok 4 Heavy rather than Grok 4 (the fast, single-agent model).

Dumb (?) question but how is Google's approach here different than Mixture of Experts? Where instead of training different experts to have different model weights you just count on temperature to provide diversity of thought. How much benefit is there in getting the diversity of thought in different runs of the same model versus running a consortium of different model weights and architectures? Is there a paper contrasting results given fixed computation between spending that compute on multiple runs of the same model vs different models?

Re: Gemini 2.5 Deep Think

#85

I would be interested in reading about how people who are paying for access to Google's top AI plan are intending to use this. Do you have any examples of immediate use-cases that might benefit? Is Google using this tool internally? One would expect them to give some examples of how it's helping internal teams accelerate or solve more challenging problems, if they were eating their own dogfood.

I'm guessing most Ultra users are there for Veo 3, where it has monetary benefits if your 3s video go viral on TikTok/Reels/Shorts

Re: Gemini 2.5 Deep Think

#86
post #76

I started doing some experimentation with this new Deep Think agent, and after five prompts I reached my daily usage limit. For $250 USD/mo that’s what you’ll be getting folks. It’s just bizarrely uncompetitive with o3-pro and Grok 4 Heavy. Anecdotally (from my experience) this was the one feature that enthusiasts in the AI community were interested in to justify the exorbitant price of Google’s Ultra subscription. I…

What was the experimentation? Can you share with us so we can see how "bizarrely uncompetitive" it is?

Bizarrely uncompetitive is referencing the 5 uses per day not the performance itself

Re: Gemini 2.5 Deep Think

#87
This comes at a time where my experience with Gemini is lacking, it seems to get worse. It's not picking up on my intention, sometimes replies in the wrong language, etc. Either that or I am just transparent that it's a tool and its feelings are hurt. I've had to call it a moron several times, and it was funny when it started reprimanding me for my foul language once. But it was wrong. This behavior seems new. I could never trust it to not do random edits everywhere in a document, so nowadays I use it to check Claude, which can be trusted with a document.

Re: Gemini 2.5 Deep Think

#88

Ladies and Gentlemen, Here's Gemini Deep Think when prompted with: "Create a svg of a pelican riding on a bicycle" https://www.svgviewer.dev/s/5R5iTexQ Beat Simon Willison to it :)

It was an expensive SVG, but it did a good job.

The bike is an actual bike with a diamond frame.

Re: Gemini 2.5 Deep Think

#89
post #69

Ladies and Gentlemen, Here's Gemini Deep Think when prompted with: "Create a svg of a pelican riding on a bicycle" https://www.svgviewer.dev/s/5R5iTexQ Beat Simon Willison to it :)

Can it do circuit diagrams? Because that's one practical area where I think the AI models are lacking.

Not yet, or schemas. It can do netlists, though! But it's much harder to go from "Netlist -> Diagram/Schema" than the other way around :(

Re: Gemini 2.5 Deep Think

#90

I started doing some experimentation with this new Deep Think agent, and after five prompts I reached my daily usage limit. For $250 USD/mo that’s what you’ll be getting folks. It’s just bizarrely uncompetitive with o3-pro and Grok 4 Heavy. Anecdotally (from my experience) this was the one feature that enthusiasts in the AI community were interested in to justify the exorbitant price of Google’s Ultra subscription. I…

> It’s just bizarrely uncompetitive with o3-pro and Grok 4 Heavy.

In my experience Grok 4 and 4 Heavy have been crap. Who cares how many requests you get with it when the response is terrible. Worst LLM money I’ve spent this year and I’ve spent a lot.

Post reply on HN