Live data from Hacker News

Introducing deep research

openai.com

351–360 of 445 posts

Re: Introducing deep research

#352

Earlier quoted context omitted.

> Can't we just send LLMs back to the drawing board until they have some semblance of reliability? Well at this point they've certainly proven a net gain for everyone regardless of the occasional nonsense they spew.

That is... debatable. You may be entirely inside the bubble, there.

Not sure if this was posted as humour, but I don't feel that way. In today's world, where I certainly would consider taking the blue pill, I'm having a blast with LLMs!

It has helped me learn stuff incredibly faster. Especially I find them useful for filling the gaps of knowledge and exploring new topics in my own way and language, without needing to wait an answer from a human (that could also be wrong).

Why does it feel, that "we are entirely inside the bubble" for you?

Re: Introducing deep research

#353
When I saw new to llms, I used Bing ai in a fun way. So when I was writing my report, it was sometimes hard to find discussions or material about a certain topic.

What I did was to ask Bing ai about that topic and it returned information aswell as sources to where it found those, so I picked up all those links and researched them myself.

Bing ai was a great resource for finding relevant links, this was until I found out about perplexity, my life haven't been the same since.

Re: Introducing deep research

#354

Earlier quoted context omitted.

> Can't we just send LLMs back to the drawing board until they have some semblance of reliability? Well at this point they've certainly proven a net gain for everyone regardless of the occasional nonsense they spew.

That is... debatable. You may be entirely inside the bubble, there.

You overestimate the importance of being correct

Re: Introducing deep research

#355
post #74
post #56

This is terrifying. Even though they acknowledge the issues with hallucinations/errors, that is going to be completely overlooked by everyone using this, and then injecting the outputs into their own powerpoints. Management Consulting was bad enough before the ability to mass produce these graphs and stats on a whim. At least there was some understanding behind the scenes of where the numbers came from, and sources w…

> At least there was some understanding behind the scenes of where the numbers came from, and sources would/could be provided. Oh Sweet summer child.

Hi tmnvdb, since you seem to love these super smart LLMs I thought it would be fun to have openais o3-mini-high analyze your recent comments in contrast to the Hacker News Comment Guidelines. Here is the output it gave me, hope it helps you:

------

Hey, I've noticed a few things in your style that are both strengths and opportunities for improvement:

Strengths:

- You clearly have deep knowledge and back up your points with solid data and examples.

- Your confidence and detailed analysis make your arguments compelling.

Opportunities:

- At times, your tone can feel a bit combative, which might shut down conversation.

- Focusing on critiquing ideas rather than questioning someone's honesty can help keep the discussion constructive.

- A clearer structure in longer posts could make your points even more accessible.

Overall, your passion and expertise shine through—tweaking the tone a bit might help foster even more productive debates.

------

Just reply here if you want the full 500+ words analysis that goes into more detail.

Re: Introducing deep research

#356

I just gave it a whirl. Pretty neat, but definitely watch out for hallucinations. For instance, I asked it to compile a report on myself (vain, I know.) In this 500-word report (ok, I'm not that important, I guess), it made at least three errors. It stated that I had 47,000 reputation points on Stack Overflow -- quite a surprise to me, given my minimal activity on Stack Overflow over the years. I popped over to the l…

Interesting!

I wonder if it’s carried over too much of that ‘helpful’ DNA from 4o’s RLHF. In that case, maybe asking for 500 words was the difficult part — it just didn’t have enough to say based on one SO post and one article, but the overall directives assume there is, and so the model is put into a place where it must publish..

Put another way, it seems this model faithfully replicates the incentives most academics have — publish a positive result, or get dinged. :)

Did it pick up your HN comments? Kadua claims that’s more than enough to roast me, … and it’s not wrong. It seems like there’s enough detail about you (or me) there to do a better job summarizing.

Re: Introducing deep research

#357
This is so lame. This feels like another desperate attempt to stay relevant cobbled together after the DeepSeek announcement last week. What was the other attempt they made? Skip a version number to seem like more progress was made (o1->o3)? From what I can tell "o3" is just the same as o1 with an extra reasoning-effort parameter.

Oh and "Deep research" is available to people on the $200 per month plan? Lol - cool. I've been using DeepSeek a lot more recently and it's so incredibly good even with all the scaling issues.

Re: Introducing deep research

#358

Earlier quoted context omitted.

"Pretty neat, but definitely watch out for hallucinations." We'd never hire someone who just makes stuff up (or at least keep them employed for long). Why are we okay with calling "AI" tools like this anything other than curious research projects? Can't we just send LLMs back to the drawing board until they have some semblance of reliability?

> Can't we just send LLMs back to the drawing board until they have some semblance of reliability? Well at this point they've certainly proven a net gain for everyone regardless of the occasional nonsense they spew.

No, from the research around it the findings are mixed. There is no consensus that it's net gain.

Re: Introducing deep research

#359
post #24

If I understood the graphs correctly, it only achieves 20% pass rate on their internal tests. So I have to wait 30min and pay a lot of money just to sift through walls of most likely incorrect text? Unless the possibility of hallucinations is negligible, this is just way too much content to review at once. The process probably needs to be a lot more iterative.

Yeah it can be more iterative. Just use individual queries and build on it yourself. This is all this is doing. It's a trick, and OpenAI is a PR hype company at this stage.

Re: Introducing deep research

#360
post #182
post #156

Earlier quoted context omitted.

No it is not an actual question on this exam. From the paper: “To ensure question quality and integrity, we enforce strict submission criteria. Questions should be precise, unambiguous, solvable, and non-searchable , ensuring models cannot rely on memorization or simple retrieval methods. All submissions must be original work or non-trivial syntheses of published information, though contributions from unpublished res…

It's example #7 on https://lastexam.ai/

This is an example of the submitted questions. Because it is possible to search it on the web, it is not an example of the accepted questions.
Post reply on HN