Live data from Hacker News

Gemini 2.5 Deep Think

blog.google

21–30 of 259 posts

Re: Gemini 2.5 Deep Think

#21
post #7

Approach is analogous to Grok 4 Heavy: use multiple "reasoning" agents in parallel and then compare answers before coming back with a single response, taking ~30 minutes. Great results, though it would be more fair for the benchmark comparisons to be against Grok 4 Heavy rather than Grok 4 (the fast, single-agent model).

I am surprised such a simple approach has taken so long to be actually used. My first image description cli attempt did basically that: Use n to get several answers and another pass to summarize.

It's an expensive approach, and depends on assessment being easy, which is often not the case.

Re: Gemini 2.5 Deep Think

#22

Approach is analogous to Grok 4 Heavy: use multiple "reasoning" agents in parallel and then compare answers before coming back with a single response, taking ~30 minutes. Great results, though it would be more fair for the benchmark comparisons to be against Grok 4 Heavy rather than Grok 4 (the fast, single-agent model).

Is o3-pro the same as these?

Re: Gemini 2.5 Deep Think

#23

I can’t even convince Gemini CLI while planning things to not go off and make a bunch of random changes on its own, even after being very clear not to do so, intercepting to tell it to stop doing that, then it just continues on fucking everything up.

Agents muddy the waters.

Claude Code gets the most out of Anthropic’s models, that’s why people love it.

Conversely, Gemini CLI makes Gemini Pro 2.5 less capable than the model itself actual is.

It’s such a stark difference I’ve given up using Gemini CLI even with it being free, but still use it for situations amenable to a prompt interface on a regular basis. It’s a very strong model.

Re: Gemini 2.5 Deep Think

#24
post #7

Approach is analogous to Grok 4 Heavy: use multiple "reasoning" agents in parallel and then compare answers before coming back with a single response, taking ~30 minutes. Great results, though it would be more fair for the benchmark comparisons to be against Grok 4 Heavy rather than Grok 4 (the fast, single-agent model).

I am surprised such a simple approach has taken so long to be actually used. My first image description cli attempt did basically that: Use n to get several answers and another pass to summarize.

People have played with (multi-) agentic frameworks for LLMs from the very beginning but it seems like only now with powerful reasoning models it is really making a difference.

Re: Gemini 2.5 Deep Think

#25

You can't go anywhere without having Gemini shoved in your face. I had an immediate visceral reaction to this.

I’m never the one to defend AI, but what do you mean? Is it the “AI overview” that pops up on Google? Other than that, I would say Gemini is definitely less in your face than ChatGPT for example

Re: Gemini 2.5 Deep Think

#27

Approach is analogous to Grok 4 Heavy: use multiple "reasoning" agents in parallel and then compare answers before coming back with a single response, taking ~30 minutes. Great results, though it would be more fair for the benchmark comparisons to be against Grok 4 Heavy rather than Grok 4 (the fast, single-agent model).

Grok-4 heavy benchmarks used tools, which trivializes a lot of problems.

Re: Gemini 2.5 Deep Think

#28

You can't go anywhere without having Gemini shoved in your face. I had an immediate visceral reaction to this.

I’m never the one to defend AI, but what do you mean? Is it the “AI overview” that pops up on Google? Other than that, I would say Gemini is definitely less in your face than ChatGPT for example

My company uses google workspace and every google doc, spreadsheet, calendar, online meeting and search puts nonstop callouts and messages about using Gemini. It's gotten so bad that I'm about to try building a browser extension to block that bullshit. It clutters the UI and nags. If I wanted that crap, I'd turn it on.
Post reply on HN