Live data from Hacker News

Introducing deep research

openai.com

341–350 of 445 posts

Re: Introducing deep research

#341

I just gave it a whirl. Pretty neat, but definitely watch out for hallucinations. For instance, I asked it to compile a report on myself (vain, I know.) In this 500-word report (ok, I'm not that important, I guess), it made at least three errors. It stated that I had 47,000 reputation points on Stack Overflow -- quite a surprise to me, given my minimal activity on Stack Overflow over the years. I popped over to the l…

Interesting

You might find it amusing to compare it to: https://hn-wrapped.kadoa.com/timabdulla

(Ref:https://news.ycombinator.com/item?id=42857604)

Re: Introducing deep research

#342

I just gave it a whirl. Pretty neat, but definitely watch out for hallucinations. For instance, I asked it to compile a report on myself (vain, I know.) In this 500-word report (ok, I'm not that important, I guess), it made at least three errors. It stated that I had 47,000 reputation points on Stack Overflow -- quite a surprise to me, given my minimal activity on Stack Overflow over the years. I popped over to the l…

> Pretty neat, but definitely watch out for hallucinations.

That would be exactly my verdict of any product based on LLMs in the past few years.

Re: Introducing deep research

#343

It is actually interesting for people working in academia. I would like to test it but no way I can afford $200/m right now. Can someone test it with this prompt. "As a research assistant with comprehensive knowledge of particle physics, please provide a detailed analysis of next-generation particle collider projects currently under consideration by the international physics community. The analysis should encompass t…

Ran it for you using o3-high! Here's a link to the results: https://chatgpt.com/share/67a0b227-8ee4-800f-a8ed-882e7bab97... Hopefully helpful, happy to test others for you :) -- Raw output -- Understood. I will conduct a detailed technical analysis of next-generation particle collider projects, focusing on the Future Circular Collider (FCC), International Linear Collider (ILC), Compact Linear Collider (CLIC), Muon Co…

Honestly, these are the smartest and overall best LLM outputs I've ever seen to date. Loving Deep Research, feels like another level up in the race

Re: Introducing deep research

#344

I just gave it a whirl. Pretty neat, but definitely watch out for hallucinations. For instance, I asked it to compile a report on myself (vain, I know.) In this 500-word report (ok, I'm not that important, I guess), it made at least three errors. It stated that I had 47,000 reputation points on Stack Overflow -- quite a surprise to me, given my minimal activity on Stack Overflow over the years. I popped over to the l…

"Pretty neat, but definitely watch out for hallucinations."

We'd never hire someone who just makes stuff up (or at least keep them employed for long). Why are we okay with calling "AI" tools like this anything other than curious research projects?

Can't we just send LLMs back to the drawing board until they have some semblance of reliability?

Re: Introducing deep research

#345

Gemini has had this for a month or two, also named "Deep Research" https://blog.google/products/gemini/google-gemini-deep-resea... Meta question: what's with all of the naming overlap in the AI world? Triton (Nvidia, OpenAI) and Gro{k,q} (X.ai, groq, OpenAI) all come to mind

> Gemini has had this for a month or two,

Would have loved to try it when they released it, but I'm apparently in the wrong country. I think it's not available outside the US (?). OpenAI and DeepSeek have no such issues. It's a bummer really, I'm happy paying for this but they don't want me to.

Re: Introducing deep research

#346

I just gave it a whirl. Pretty neat, but definitely watch out for hallucinations. For instance, I asked it to compile a report on myself (vain, I know.) In this 500-word report (ok, I'm not that important, I guess), it made at least three errors. It stated that I had 47,000 reputation points on Stack Overflow -- quite a surprise to me, given my minimal activity on Stack Overflow over the years. I popped over to the l…

"Pretty neat, but definitely watch out for hallucinations." We'd never hire someone who just makes stuff up (or at least keep them employed for long). Why are we okay with calling "AI" tools like this anything other than curious research projects? Can't we just send LLMs back to the drawing board until they have some semblance of reliability?

> Can't we just send LLMs back to the drawing board until they have some semblance of reliability?

Well at this point they've certainly proven a net gain for everyone regardless of the occasional nonsense they spew.

Re: Introducing deep research

#347
post #130

Earlier quoted context omitted.

I really don't like the snarky tone of the parent comment. Nonetheless, I don't think this is even something that can easily be benchmarked. I'd recommend you take a look at aider [1], and consider how I drew similarities between it and what's presented here. Has ClosedAI presented any benchmarks / evaluation protocols? [1] https://aider.chat/

Yes, they show benchmarks in the article linked here. Did you not read it?

I don’t think you actually read it. The benchmarks are in reference to the model that’s underlying deep-research, and not deep-research itself. For the latter, they have anecdata from scientists.

Re: Introducing deep research

#348

Earlier quoted context omitted.

"Pretty neat, but definitely watch out for hallucinations." We'd never hire someone who just makes stuff up (or at least keep them employed for long). Why are we okay with calling "AI" tools like this anything other than curious research projects? Can't we just send LLMs back to the drawing board until they have some semblance of reliability?

> Can't we just send LLMs back to the drawing board until they have some semblance of reliability? Well at this point they've certainly proven a net gain for everyone regardless of the occasional nonsense they spew.

That is... debatable. You may be entirely inside the bubble, there.

Re: Introducing deep research

#349
post #341

I just gave it a whirl. Pretty neat, but definitely watch out for hallucinations. For instance, I asked it to compile a report on myself (vain, I know.) In this 500-word report (ok, I'm not that important, I guess), it made at least three errors. It stated that I had 47,000 reputation points on Stack Overflow -- quite a surprise to me, given my minimal activity on Stack Overflow over the years. I popped over to the l…

Interesting You might find it amusing to compare it to: https://hn-wrapped.kadoa.com/timabdulla (Ref: https://news.ycombinator.com/item?id=42857604 )

This is... very uncomfortable. An (expanded) AI summary of my HN and reddit usage would appear to be a pretty complete representation of my "online" identity/character. I remember when people would browse your entire comment history just to find something to discredit you on reddit, and that behavior was _heavily_ discouraged. Now, we can just run an AI model to follow you and sentence you to a hell of being permanently discredited online. Give it a bunch of accounts to rotate through, send some voting power behind it (reddit or hn), and just pick apart every value you hold. You could obliterate someone's will to discuss anything online. You could effectively silence all but the most stubborn, and those people you would probably drive insane.

It's a very interesting usecase though, filter through billions of comments and give everyone a score on which real life person they probably are. I wonder if say, Ted Cruz hides behind a username somewhere.

Re: Introducing deep research

#350

I just gave it a whirl. Pretty neat, but definitely watch out for hallucinations. For instance, I asked it to compile a report on myself (vain, I know.) In this 500-word report (ok, I'm not that important, I guess), it made at least three errors. It stated that I had 47,000 reputation points on Stack Overflow -- quite a surprise to me, given my minimal activity on Stack Overflow over the years. I popped over to the l…

"Pretty neat, but definitely watch out for hallucinations." We'd never hire someone who just makes stuff up (or at least keep them employed for long). Why are we okay with calling "AI" tools like this anything other than curious research projects? Can't we just send LLMs back to the drawing board until they have some semblance of reliability?

[deleted]
Post reply on HN