Live data from Hacker News

Introducing deep research

openai.com

311–320 of 445 posts

Re: Introducing deep research

#311
post #271

Earlier quoted context omitted.

Traditional ML is no stranger to measuring accuracy in terms of agreement with human evaluators.

A customer churn model or revenue forecast did have hard objective data (ground truth) to compare against - isn't it?

you seem to think all classical ML models were supervised, which isn't true. and we have metrics for unsupervised approaches as well

Re: Introducing deep research

#312

For “deep research” I’m also reading “getting the answers right”. Most people I talk to are at the point now where getting completely incorrect answers 10% of the time — either obviously wrong from common sense, or because the answers are self contradictory — undermines a lot of trust in any kind of interaction. Other than double checking something you already know, language models aren’t large enough to actually kno…

> and also faster than it takes me to verify the answer given by the machine.

I always thought there was a kind of NP-flavor to the problems for which LLMs-like AI are helpful in practice, in the sense that solving the problem may be hard but checking the solution must be fast.

Unless the domain can accomodate errors/hallucination, checking the solution (by a human) should be exponentially faster than finding it (by some AI) otherwise there's little practical gain.

Re: Introducing deep research

#313
Late Sunday night, I gained access to OpenAI’s newly launched Deep Research and immediately tested it on a draft blog post about Uniform Electronic Transactions Act (UETA) compliance and AI-agent error handling [1]. Here’s what I found:

Within minutes, it generated a detailed, well-cited research report that significantly expanded my original analysis, covering: * Legal precedents & case law interpretations (including a nuanced breakdown of UETA Section 10). * Comparative international frameworks (EU, UK, Canada). * Real-world technical implementations (Stripe’s AI-driven transaction handling). * Industry perspectives & business impact (trust, risk allocation, compliance). * Emerging regulatory standards (EU AI Act, FTC oversight, ISO/NIST AI governance).

What stood out most was its ability to: - Synthesize complex legal, business, and technical concepts into clear, actionable insights. - Connect legal frameworks, industry trends, and real-world case studies. - Maintain a business-first focus, emphasizing practical benefits. - Integrate 2024 developments with historical context for a deeper analysis.

The depth and coherence of the output were comparable to what I would expect from a team of domain experts—but delivered in a fraction of the time.

From the announcement: Deep Research leverages OpenAI’s next-generation model, optimized for multi-step research, reasoning, and synthesis. It has already set new performance benchmarks, achieving 26.6% accuracy on Humanity’s Last Exam (the highest of any OpenAI model) and a 72.57% average accuracy on the GAIA Benchmark, demonstrating advanced reasoning and research capabilities.

Currently available to Pro users (with up to 100 queries per month), it will soon expand to Plus and Team users. While OpenAI acknowledges limitations—such as occasional hallucinations and challenges in source verification—its iterative deployment strategy and continuous refinement approach are promising.

My key takeaway: This LLM agent-based tool has the potential to save hours of manual research while delivering high-quality, well-documented outputs. Automating tasks that traditionally require expert-level investigation, it can complete complex research in 5–30 minutes (just 6 minutes for my task), with citations and structured reasoning.

I don’t see any other comments yet from people who have actually used it, but it’s only been a few hours.I’d love to hear how it’s performing for others. What use cases have you explored? How did it do?

(Note: This review is based on a single use case. I’ll provide further updates as I conduct broader testing.)

[1] https://www.dazzagreenwood.com/p/ueta-and-llm-agents-a-deep-...

Re: Introducing deep research

#314

Earlier quoted context omitted.

Why doesn't the iPhone screen spammers yet? Pixel has had this feature for a decade.

Pixel hasn’t even been around for a decade.

The Pixel branding is 12 years old, and IIRC this feature also existed in Nexus before that.

Re: Introducing deep research

#315
This make sense, I often use the normal search feature to research a very large ammount of information and it mostly does not work well. If the new search feature increases the number of websites scrapped and the pertinence of the websites, I'm all in.

Re: Introducing deep research

#316
I just gave it a whirl. Pretty neat, but definitely watch out for hallucinations. For instance, I asked it to compile a report on myself (vain, I know.) In this 500-word report (ok, I'm not that important, I guess), it made at least three errors.

It stated that I had 47,000 reputation points on Stack Overflow -- quite a surprise to me, given my minimal activity on Stack Overflow over the years. I popped over to the link it had cited (my profile on Stack Overflow) and it seems it confused my number of people reached (47k) with my reputation, a sadly paltry 525.

Then it cited an answer I gave on Stack Overflow on the topic of monkey-patching in PHP, using this as evidence for my technical expertise. Turns out that about 15 years ago, I _asked_ a question on this topic, but the answer was submitted by someone else. Looks like I don't have much expertise, after all.

Finally, it found a gem of a quote from an interview I gave. Or wait, that was my brother! Confusingly, we founded a company together, and we were both mentioned in the same article, but he was the interviewee, not I.

I would say it's decent enough for a springboard, but you should definitely treat the output with caution and follow the links provided to make sure everything is accurate.

Re: Introducing deep research

#317

For “deep research” I’m also reading “getting the answers right”. Most people I talk to are at the point now where getting completely incorrect answers 10% of the time — either obviously wrong from common sense, or because the answers are self contradictory — undermines a lot of trust in any kind of interaction. Other than double checking something you already know, language models aren’t large enough to actually kno…

> Most people I talk to are at the point now where getting completely incorrect answers 10% of the time

A year back that number was 30%, and a couple of years back it was 60%. There will be a point where it'll be good enough. There are also better and better ways to verify answers these days.

It'll never be a solution for everything, but that's similar to many engineering problems we have: for example, ORMs aren't great for all types of queries, but they're sufficient for a good part of them.

Re: Introducing deep research

#318

Late Sunday night, I gained access to OpenAI’s newly launched Deep Research and immediately tested it on a draft blog post about Uniform Electronic Transactions Act (UETA) compliance and AI-agent error handling [1]. Here’s what I found: Within minutes, it generated a detailed, well-cited research report that significantly expanded my original analysis, covering: * Legal precedents & case law interpretations (includin…

I tried it on a few things I was familiar with just to assess its reliability.

The first was on a topic with which I am deeply familiar -- myself -- and it made three factual errors in a 500-word report: https://news.ycombinator.com/item?id=42916899

The second was a task to do an industry analysis on a space in which I worked for about ten years. I think its overall synthesis was good (it accorded with my understanding of the space), but there were a number of errors in the statistics and supporting evidence it compiled, based upon my random review of the source material.

I think the product is cool and will definitely be helpful, but I would still recommend verifying its outputs. I think the process of verification is less time-consuming than the process of researching and writing, so that is likely an acceptable compromise in many cases.

Re: Introducing deep research

#319

    can sometimes hallucinate facts in responses or make incorrect inferences, though at a notably lower rate than existing ChatGPT models, according to internal evaluations. It may struggle with distinguishing authoritative information from rumors, and currently shows weakness in confidence calibration, often failing to convey uncertainty accurately
Taken from the limitations section.

These tools are just good at creating pollution. I don't see the point of delegating a (not just) research where 1% blatant mistakes are acceptable. These need much better grounding before handing out to masses.

I can not take any output by these tools (google summaries, comment summaries by amazon, youtube summaries etc etc) while knowing for a fact some of that is a total lie. I can not tell which part is a lie. e.g. If LLM says that in any given text the sentiment is divided, it could be just one person with an opposing view.

If same task was given to a person, I could reason with that person on any conclusion. These tools will reason on their hallucinations.

Re: Introducing deep research

#320
There is no way I'll read all that text from the demos...

AskPandi has a similar feature called "Super Search" that essentially checks more sources and self validates it's own answers.

iT's AgEnTic.

The answers are easier to digest, if you search for products, you'll get a list of products with images, prices and retailers.

Post reply on HN