Deep research used to mean spending a weekend with grep and a coffee pot. Now it’s just autocomplete with a confidence interval.
Deep Research is now available on Gemini 2.5 Pro Experimental
11–20 of 26 posts
Re: Deep Research is now available on Gemini 2.5 Pro Experimental
#12Earlier quoted context omitted.
Just did a test last week and OpenAIs research was way better. Found 10x more sources and did an overall pretty great job The task was to lookup information about a late distant family member who had been a prominent employee in a certain foreign government about 100 years ago Gemini barely scratched the surface and pretty much gave up ChatGPT on the other hand, kept building up on its research, connecting the dots a…
Would love to see this repeated with this latest version from Google. Man, what's really missing from all of this is a 3rd party AI Consumer Reports type site for all of these LLM tools. Whoever does this thing that does not scale will have a highly referenced site on their hands.
Re: Deep Research is now available on Gemini 2.5 Pro Experimental
#13Earlier quoted context omitted.
Just did a test last week and OpenAIs research was way better. Found 10x more sources and did an overall pretty great job The task was to lookup information about a late distant family member who had been a prominent employee in a certain foreign government about 100 years ago Gemini barely scratched the surface and pretty much gave up ChatGPT on the other hand, kept building up on its research, connecting the dots a…
Would love to see this repeated with this latest version from Google. Man, what's really missing from all of this is a 3rd party AI Consumer Reports type site for all of these LLM tools. Whoever does this thing that does not scale will have a highly referenced site on their hands.
Re: Deep Research is now available on Gemini 2.5 Pro Experimental
#14Has anyone tested googles functionality vs ChatGPT? I have lightly played around with it but felt that generally ChatGPTs implementation was a little more educated sounding and felt like it took whatever necessary persona well.
My ranking openai > grok 3 deeper > Gemini 2.0 pro. All have been terrible for the 100 or so times I’ve used them (all SWE / finance related in some way)
Re: Deep Research is now available on Gemini 2.5 Pro Experimental
#15Re: Deep Research is now available on Gemini 2.5 Pro Experimental
#16> In our testing, raters preferred the reports generated by Gemini Deep Research powered by 2.5 Pro over other leading deep research providers by more than a 2-to-1 margin. Are these raters experts in the field the report was written on? Did they rate the reports on factuality, broadness, and insights? These sort of tests (and RLHF in general) are the reason that LLMs often respond with "Great question, you are exact…
OpenAI's Deep Research seems oddly restricted in the number of sources it uses, eg repeating one survey article over and over. I suspect it is just too draining and demoralizing for RLHFers to check Deep Research's citations (especially without a formal bibliography).
Re: Deep Research is now available on Gemini 2.5 Pro Experimental
#17Earlier quoted context omitted.
Would love to see this repeated with this latest version from Google. Man, what's really missing from all of this is a 3rd party AI Consumer Reports type site for all of these LLM tools. Whoever does this thing that does not scale will have a highly referenced site on their hands.
Throughout the entire 20th century the main determinant of a Consumer Reports rating for a car was whether you could put a wheelchair in the trunk. Hopefully the AI agent industry does not sprout a similarly worthless metric.
Re: Deep Research is now available on Gemini 2.5 Pro Experimental
#18Has anyone tested googles functionality vs ChatGPT? I have lightly played around with it but felt that generally ChatGPTs implementation was a little more educated sounding and felt like it took whatever necessary persona well.
I haven’t used 2.5 pro just 2.0 pro. It was inferior to OpenAI (which isn’t that good). My ranking openai > grok 3 deeper > Gemini 2.0 pro. All have been terrible for the 100 or so times I’ve used them (all SWE / finance related in some way)
Re: Deep Research is now available on Gemini 2.5 Pro Experimental
#19Earlier quoted context omitted.
Would love to see this repeated with this latest version from Google. Man, what's really missing from all of this is a 3rd party AI Consumer Reports type site for all of these LLM tools. Whoever does this thing that does not scale will have a highly referenced site on their hands.
Isn't that what llmarena does?
I want a staff of human testers, each with domain expertise. If the goal is to replace humans, should there not be a real human metric?
I want a physicist asking their battery of physics questions, 4 different kinds of devs asking their battery of dev problems, a couple chefs asking for cooking techniques, etc.
Now on to "Deep Research," 6 different kinds of OSINT/secondary analysts who ask new problems each time, and compare it to their days of human work.
We really need this as a species, otherwise the brain dead C-Suites of the world are going to keep buying the hype, which is often very premature. This could have real consequences, and it apparently already has.
It's insane to me that we are investing, what, almost $1T into LLMs, and have not spent the ~$1.5M/year to do what I described above.