Pro user. No access like everyone else. OpenAI is very much in an existential crisis and their poor execution is not helping their cause. Operator or “deep research” should be able to assume the role of a Pro user, run a quick test, and reliably report on whether this is working before the press release right?
Introducing deep research
231–240 of 445 posts
Re: Introducing deep research
#232I do use these systems from time to time, but it just never renders any specific information that would make it great research.
Re: Introducing deep research
#233Is this ability really a prerequisite to AGI and ASI? Reasoning, problem solving, research validation - at the fundamental outset it is all refinement thinking. Research is one of those areas where I remain skeptical it is that important because the only valid proof is in the execution outcome, not the compiled answer. For instance you can research all you want about the best vacuum on the internet but until you try…
So you wouldn't use this tool for those types of use cases.
But still, a valid point. I recall I once wanted to compare Hydroflask, Klean Kanteen and Thermos to see how they perform for hot/cold drinks. I was looking specifically for articles/posts where people had performed actual measurements. But those were very hard to find, with almost all Google hits being generic comparisons with no hard data. That didn't stop them from ranking ("Hydroflask is better for warm drinks!")
Would I be able to get this to ignore all of those and use only ones where actual experiments were performed. And moreover, filter out duplicates (e.g. one guy does an experiment, and several other bloggers link to his post and repeat his findings in their own posts - it's one experiment but with many search results).
Re: Introducing deep research
#234If I understood the graphs correctly, it only achieves 20% pass rate on their internal tests. So I have to wait 30min and pay a lot of money just to sift through walls of most likely incorrect text? Unless the possibility of hallucinations is negligible, this is just way too much content to review at once. The process probably needs to be a lot more iterative.
Here's an example of the type of question it is acheiving 20% on; The set of natural transformations between two functors F,G :C→DF,G:C→D can be expressed as the end Nat(F,G)≅∫AHomD(F(A),G(A)). Nat(F,G)≅∫A HomD (F(A),G(A)). Define set of natural cotransformations from FF to GG to be the coend CoNat(F,G)≅∫AHomD(F(A),G(A)). CoNat(F,G)≅∫AHomD (F(A),G(A)). Let: - F=B∙(Σ4)∗/F=B∙ (Σ4 )∗/ be the under ∞∞-category of the ne…
Re: Introducing deep research
#235Earlier quoted context omitted.
> but this is being powered by the unreleased o3 model What makes you believe that?
they explicitly stated it in the launch
> Powered by a version of the upcoming OpenAI o3 model that’s optimized for web browsing and data analysis, it leverages reasoning to search, interpret, and analyze massive amounts of text, images, and PDFs on the internet, pivoting as needed in reaction to information it encounters.
If that's what you're referring to, then it doesn't seem that "explicit" to me. For example, how do we know that it doesn't use less thinking than o3-mini? Google's version of deep research uses their "not cutting edge version" 1.5 model, after all. Are you referring to something else?
Re: Introducing deep research
#236Earlier quoted context omitted.
I'm sure o3 will be a generation ahead of whatever deepseek, google and meta are doing today when it launches in 10 months, super impressive stuff.
I’m not sure if you’re implying this subtly in your comment or not, as it’s early here, but it does of course need to be a generation ahead of what 10 months of their competitors moving forward have done too. Nobody is standing still
Re: Introducing deep research
#237Earlier quoted context omitted.
Only if you are asking questions at the level of a cutting edge benchmark
This is one of the actual questions: > In Greek mythology, who was Jason's maternal great-grandfather? https://www.google.com/search?q=In+Greek+mythology%2C+who+wa...
Re: Introducing deep research
#238Earlier quoted context omitted.
It's more likely this is a response to Gemini Deep Research released in December https://blog.google/products/gemini/google-gemini-deep-resea...
That Google product isnt that good, it can't really replace research done by a person.
Re: Introducing deep research
#239I’m a researcher and honestly not worried. 1. Developing the right question has always been the largest barrier to great research. Not sure OpenAI can develop the right question without the Human experience. The second biggest part of my role is influencing people that my questions are the right questions. Which is made easier when you have a thorough understanding of the first. That being said, I’m sure there will b…
These systems serve best at augmenting information discovery. When I'm tackling a new area or looking for the right terminology, these models provide a quick shortcut because they have good probabilistic "understanding" of my naive, jargon-free description. This allows me to pull in all of the jargon for the area of research I'm interested in, and move on to actually useful resources, whether that be journal articles, textbooks, or - rarely - online posts/blogs/videos.
the current "meta" is probably something like Elicit + notebookLM + Claude for accelerating understanding of complex topics and extracting useful parts. But, again, each step requires that I am closely involved, from selecting the "correct" papers, to carefully aggregating and grooming the information pulled in from notebookLM, to judging the the usefulness of Claude's attempts to extract what I have asked for
Re: Introducing deep research
#240Can someone test it with this prompt.
"As a research assistant with comprehensive knowledge of particle physics, please provide a detailed analysis of next-generation particle collider projects currently under consideration by the international physics community.
The analysis should encompass the major proposed projects, including the Future Circular Collider (FCC) at CERN, International Linear Collider (ILC), Compact Linear Collider (CLIC), various Muon Collider proposals, and any other significant projects as of 2024.
For each proposal, examine the planned energy ranges and collision types, estimated timeline for construction and operation, technical advantages and challenges, approximate costs, and key physics goals. Include information about current technical design reports, feasibility studies, and the level of international support and collaboration.
Present a thorough comparative analysis that addresses technical feasibility, cost-benefit considerations, scientific potential for new physics discoveries, timeline to first data collection, infrastructure requirements, and environmental impact. The projects should be compared in terms of their relative strengths, weaknesses, and potential contributions to advancing our understanding of fundamental physics.
Please format the response as a structured technical summary suitable for presentation at a topical meeting of particle physicists. Where appropriate, incorporate relevant figures and tables to facilitate clear comparisons between proposals. Base your analysis on information from peer-reviewed sources and official design reports, focusing on the most current available data and design specifications.
Consider the long-term implications of each proposal, including potential upgrade paths, flexibility for future modifications, and integration with existing research infrastructure."