Live data from Hacker News

Introducing deep research

openai.com

171–180 of 445 posts

Re: Introducing deep research

#171

Not sure if people picked up on it, but this is being powered by the unreleased o3 model. Which might explain why it leaps ahead in benchmarks considerably and aligns with the claims o3 is too expensive to release publicly. Seems to be quite an impressive model and the leading out of Google, DeepSeek and Perplexity.

> Which might explain why it leaps ahead in benchmarks considerably and aligns with the claims o3 is too expensive to release publicly

It's the only tool/system (I won't call it an LLM) in their released benchmarks that has access to tools and the web. So, I'd wager the performance gains are strictly due to that.

If an LLM (o3) is too expensive to be released to the public, why would you use it in a tool that has to make hundreds of inference calls to it to answer a single question? You'd use a much cheaper model. Most likely o3-mini or o1-mini combined with o4-mini for some tasks.

Re: Introducing deep research

#172
post #155
post #119

Earlier quoted context omitted.

Because maybe you want to, but you have a boss breathing down your neck and KPIs to meet and you haven't slept properly in days and just need a win, so you get the AI to put together some impressive looking graphs and stats that will look impressive in that client showcase thats due in a few hours. Things aren't quite so black and white in reality.

I mean those same conditions already just lead the human to cutting corners and making stuff up themselves. You're describing the problem where bad incentives/conditions lead to sloppy work, that happens with or without AI Catching errors/validating work is obviously a different process when they're coming from an AI vs a human, but I don't see how it's fundamentally that different here. If the outputs are heavily ci…

Yep, I agree with this to some extent, but I think the difference in the future is all that stress will be bypassed and people will reach for the AI from the start.

Previously there was alot of stress/pressure which might or might not have led to sloppy work (some consultants are of a high quality). With this, there will be no stress which will (always?) lead to sloppy work. Perhaps there's an argument for the high quality consultants using the tools to produce accurate and high quality work. There will obviously be a sliding scale here. Time will tell.

I'd wager the end result will be sloppy work, at scale :-)

Re: Introducing deep research

#173
post #24

If I understood the graphs correctly, it only achieves 20% pass rate on their internal tests. So I have to wait 30min and pay a lot of money just to sift through walls of most likely incorrect text? Unless the possibility of hallucinations is negligible, this is just way too much content to review at once. The process probably needs to be a lot more iterative.

Here's an example of the type of question it is acheiving 20% on; The set of natural transformations between two functors F,G ⁣:C→DF,G:C→D can be expressed as the end Nat(F,G)≅∫AHomD(F(A),G(A)). Nat(F,G)≅∫A HomD (F(A),G(A)). Define set of natural cotransformations from FF to GG to be the coend CoNat(F,G)≅∫AHomD(F(A),G(A)). CoNat(F,G)≅∫AHomD (F(A),G(A)). Let: - F=B∙(Σ4)∗/F=B∙ (Σ4 )∗/ be the under ∞∞-category of the ne…

[dead]

Re: Introducing deep research

#174
Actually sounds pretty cool, but the graph on expert level tasks is confusing my expectations. Saying it has a pass rate of less than 20% sounds a lot like saying this thing is wrong most of the time.

Granted, these strike me as difficult tasks and I’d likely ask it to do far simpler things, but I’m not really sure what to expect from looking at these graphs.

Ah, but the fact that it bothers to cite its sources is a huge plus. Between that and its search abilities it sounds valuable to me

Re: Introducing deep research

#176
I see lots of warranted skepticism about the capabilities of this tool, but the reality is that this is an incremental step toward full automation of white collar labor. No, it will not make all analysts jobless overnight. But it may reduce hiring of said people by 5 or 10 percent. And as people get better at using the tool and the tool itself gets better, those numbers will grow. Remember that it took decades for the giant pool of typing secretaries in Mad Men to disappear, but they did disappear. Gone forever. Interestingly, anger about the diminishment of secretarial male white collar work in Germany due to the spread of the typewriter a few decades earlier was one of the drivers of the Nazi Party’s popularity (see Evans, the Rise of the Third Reich).

AI’s triumph in the white collar workplace will be gradual, not instantaneous. And it will be grimly quiet, because no one likes white collar workers the way they like blue collar workers, for some odd reason, and there’s no tradition of solidarity among white collar workers. Everyone will just look up one day and find that the local Big Corp headquarters is…empty.

Re: Introducing deep research

#177
post #94

Earlier quoted context omitted.

> they are probabilistic language models This is like arguing an Airbus cannot possibly fly because it is 165 tonnes of aluminum, steel and plastic. The proof is in the fact that it flies, not what it is constructed from.

> The proof is in the fact that it flies, not what it is constructed from. And LLMs do not. > "But it looks like reasoning to me" My condolences. You should go see a doctor about your inability to count the number of 'R's in a word.

[dead]

Re: Introducing deep research

#179

Not sure if people picked up on it, but this is being powered by the unreleased o3 model. Which might explain why it leaps ahead in benchmarks considerably and aligns with the claims o3 is too expensive to release publicly. Seems to be quite an impressive model and the leading out of Google, DeepSeek and Perplexity.

I'm sure o3 will be a generation ahead of whatever deepseek, google and meta are doing today when it launches in 10 months, super impressive stuff.

I’m not sure if you’re implying this subtly in your comment or not, as it’s early here, but it does of course need to be a generation ahead of what 10 months of their competitors moving forward have done too. Nobody is standing still

Re: Introducing deep research

#180

Earlier quoted context omitted.

I’m not sure I understand what you mean by “the button”. If you’re comparing this to DeepSeek’s copying, it’s not really the same thing right? DeepSeek essentially stole intellectual property by violating OpenAI’s terms of service. As I understand it, this is a copy of Google’s Deep Research

I chuckle every time I see this. Poor OpenAI. Meanwhile, their entire training corpus was the result of scraping the intellectual property and copyrighted materials of THE ENTIRE PUBLIC INTERNET. Woe is them to be sure.

OpenAI’s scraping will likely be ruled as fair use.
Post reply on HN