Approach is analogous to Grok 4 Heavy: use multiple "reasoning" agents in parallel and then compare answers before coming back with a single response, taking ~30 minutes. Great results, though it would be more fair for the benchmark comparisons to be against Grok 4 Heavy rather than Grok 4 (the fast, single-agent model).
I am surprised such a simple approach has taken so long to be actually used. My first image description cli attempt did basically that: Use n to get several answers and another pass to summarize.
Gemini 2.5 Deep Think
21–30 of 259 posts
Re: Gemini 2.5 Deep Think
#22Approach is analogous to Grok 4 Heavy: use multiple "reasoning" agents in parallel and then compare answers before coming back with a single response, taking ~30 minutes. Great results, though it would be more fair for the benchmark comparisons to be against Grok 4 Heavy rather than Grok 4 (the fast, single-agent model).
Re: Gemini 2.5 Deep Think
#23I can’t even convince Gemini CLI while planning things to not go off and make a bunch of random changes on its own, even after being very clear not to do so, intercepting to tell it to stop doing that, then it just continues on fucking everything up.
Claude Code gets the most out of Anthropic’s models, that’s why people love it.
Conversely, Gemini CLI makes Gemini Pro 2.5 less capable than the model itself actual is.
It’s such a stark difference I’ve given up using Gemini CLI even with it being free, but still use it for situations amenable to a prompt interface on a regular basis. It’s a very strong model.
Re: Gemini 2.5 Deep Think
#24Approach is analogous to Grok 4 Heavy: use multiple "reasoning" agents in parallel and then compare answers before coming back with a single response, taking ~30 minutes. Great results, though it would be more fair for the benchmark comparisons to be against Grok 4 Heavy rather than Grok 4 (the fast, single-agent model).
I am surprised such a simple approach has taken so long to be actually used. My first image description cli attempt did basically that: Use n to get several answers and another pass to summarize.
Re: Gemini 2.5 Deep Think
#25You can't go anywhere without having Gemini shoved in your face. I had an immediate visceral reaction to this.
Re: Gemini 2.5 Deep Think
#26Re: Gemini 2.5 Deep Think
#27Approach is analogous to Grok 4 Heavy: use multiple "reasoning" agents in parallel and then compare answers before coming back with a single response, taking ~30 minutes. Great results, though it would be more fair for the benchmark comparisons to be against Grok 4 Heavy rather than Grok 4 (the fast, single-agent model).
Re: Gemini 2.5 Deep Think
#28You can't go anywhere without having Gemini shoved in your face. I had an immediate visceral reaction to this.
I’m never the one to defend AI, but what do you mean? Is it the “AI overview” that pops up on Google? Other than that, I would say Gemini is definitely less in your face than ChatGPT for example
Re: Gemini 2.5 Deep Think
#29Available via API?
Re: Gemini 2.5 Deep Think
#30It's not yet available via an API.