Live data from Hacker News

Introducing deep research

openai.com

261–270 of 445 posts

Re: Introducing deep research

#261
post #24

If I understood the graphs correctly, it only achieves 20% pass rate on their internal tests. So I have to wait 30min and pay a lot of money just to sift through walls of most likely incorrect text? Unless the possibility of hallucinations is negligible, this is just way too much content to review at once. The process probably needs to be a lot more iterative.

Here's an example of the type of question it is acheiving 20% on; The set of natural transformations between two functors F,G ⁣:C→DF,G:C→D can be expressed as the end Nat(F,G)≅∫AHomD(F(A),G(A)). Nat(F,G)≅∫A HomD (F(A),G(A)). Define set of natural cotransformations from FF to GG to be the coend CoNat(F,G)≅∫AHomD(F(A),G(A)). CoNat(F,G)≅∫AHomD (F(A),G(A)). Let: - F=B∙(Σ4)∗/F=B∙ (Σ4 )∗/ be the under ∞∞-category of the ne…

As someone who doesn't understand anything beyond the word 'set' in that question, can anyone give an indication of how hard of a problem that actually is (within that domain)?

Also I'm curious as to what percentage of the questions in this benchmark are of this type / difficulty, vs the seemingly much easier example of "In Greek mythology, who was Jason's maternal great-grandfather?".

I'd imagine the latter is much easier for an LLM, and almost trivial for any LLM with access to external sources (such as deep research).

Re: Introducing deep research

#262
I remember about 10-15 years ago that Ray Kurzweil (who still works at Google) or someone at Google had this idea for what Google should be able to do: About doing deep research by itself with a simple search query. I can't find the source. Obviously it didn't pan out without transformers.

Re: Introducing deep research

#263
post #230

Earlier quoted context omitted.

I have pro, in US, not seeing yet

what about a full refresh of the page or perhaps jump into the dev tools and check "disable cache" could also be aggressive caching from cloudflare. could be they're just trying to announce more stuff to maintain cachet and can't yet support all users forking over 200/month.

I relogged, disabled cache and reloaded the page with Ctrl+Shift+R but it doesn't show up.

Re: Introducing deep research

#264
post #56

This is terrifying. Even though they acknowledge the issues with hallucinations/errors, that is going to be completely overlooked by everyone using this, and then injecting the outputs into their own powerpoints. Management Consulting was bad enough before the ability to mass produce these graphs and stats on a whim. At least there was some understanding behind the scenes of where the numbers came from, and sources w…

Then the hallucinated research is published in an article which is then cited by other AI research, continuing the push the false information until it’s hard to know where the lie started.

Re: Introducing deep research

#266
post #205

Earlier quoted context omitted.

When things are easy, you’re going to take the easy path even if it means quality goes down. It’s about trade offs. If you had to do it yourself, perhaps quality would have been higher because you had no other choice. Lots of kids don’t want to do homework. That said, previously many would because there wasn’t another choice. But now they can just ask ChatGPT for the answers they’ll write that down verbatim with zero…

"Lots of kids don’t want to do homework" Sure, but if you're a professional you have to care about your reputation. Presenting hallucinated cases from ChatGPT didn't go very well for that lawyer: https://www.nytimes.com/2023/05/27/nyregion/avianca-airline-...

That's a lawyer in an adverserial situation. Business consultants tell their clients what they want to believe, the facts be dammed.

Re: Introducing deep research

#267

I’m a researcher and honestly not worried. 1. Developing the right question has always been the largest barrier to great research. Not sure OpenAI can develop the right question without the Human experience. The second biggest part of my role is influencing people that my questions are the right questions. Which is made easier when you have a thorough understanding of the first. That being said, I’m sure there will b…

[deleted]

Re: Introducing deep research

#268
post #230

Earlier quoted context omitted.

what about a full refresh of the page or perhaps jump into the dev tools and check "disable cache" could also be aggressive caching from cloudflare. could be they're just trying to announce more stuff to maintain cachet and can't yet support all users forking over 200/month.

I relogged, disabled cache and reloaded the page with Ctrl+Shift+R but it doesn't show up.

same here. pro in the US and still no access. i even logged in using my phone and a different browser

Re: Introducing deep research

#269
I think deep research as a service could be a really strong use case for enterprises, as long as they have access to non-public data. I assume that most of this guarded data is high quality, and seeing progress in these areas might end up being even more impressive than it is now.

Re: Introducing deep research

#270
post #205

Earlier quoted context omitted.

"Lots of kids don’t want to do homework" Sure, but if you're a professional you have to care about your reputation. Presenting hallucinated cases from ChatGPT didn't go very well for that lawyer: https://www.nytimes.com/2023/05/27/nyregion/avianca-airline-...

That's a lawyer in an adverserial situation. Business consultants tell their clients what they want to believe, the facts be dammed.

it sounds like ai doesn't really change that situation
Post reply on HN