Live data from Hacker News

Introducing deep research

openai.com

401–410 of 445 posts

Re: Introducing deep research

#401

I just gave it a whirl. Pretty neat, but definitely watch out for hallucinations. For instance, I asked it to compile a report on myself (vain, I know.) In this 500-word report (ok, I'm not that important, I guess), it made at least three errors. It stated that I had 47,000 reputation points on Stack Overflow -- quite a surprise to me, given my minimal activity on Stack Overflow over the years. I popped over to the l…

"Pretty neat, but definitely watch out for hallucinations." We'd never hire someone who just makes stuff up (or at least keep them employed for long). Why are we okay with calling "AI" tools like this anything other than curious research projects? Can't we just send LLMs back to the drawing board until they have some semblance of reliability?

Why not just verify the output? It’s faster than generating the entire thing yourself. Why do you need perfection in a productivity tool?

Re: Introducing deep research

#402

Earlier quoted context omitted.

Ran it for you using o3-high! Here's a link to the results: https://chatgpt.com/share/67a0b227-8ee4-800f-a8ed-882e7bab97... Hopefully helpful, happy to test others for you :) -- Raw output -- Understood. I will conduct a detailed technical analysis of next-generation particle collider projects, focusing on the Future Circular Collider (FCC), International Linear Collider (ILC), Compact Linear Collider (CLIC), Muon Co…

Honestly, these are the smartest and overall best LLM outputs I've ever seen to date. Loving Deep Research, feels like another level up in the race

Thank you very much for doing that. It is actually somehow impressive. It got a lot of big picture comparison and points correct. There are problem with some details but overall it does save some work for initial search process.

What I like is that it asked you before clarifying questions before but I wonder if it just generic. Because the prompt mentioned that this would be for "presentation at a topical meeting of particle physicists" but still asked its last question about

> Intended Audience: Should the analysis assume a general physics audience or a more specialized group of particle physicists?

Also probably expected but it didn't include or reference graphs/plots.

Re: Introducing deep research

#403

I just gave it a whirl. Pretty neat, but definitely watch out for hallucinations. For instance, I asked it to compile a report on myself (vain, I know.) In this 500-word report (ok, I'm not that important, I guess), it made at least three errors. It stated that I had 47,000 reputation points on Stack Overflow -- quite a surprise to me, given my minimal activity on Stack Overflow over the years. I popped over to the l…

This is very bearish for current AI. Seems like 99% reliability is still too small with compounding errors. But I wonder of this is inherently specific to longer context or if this just depends on how it’s trained. In theory longer context => more errors

Although I think people are the same, too big problem and you are getting lost unless taking it in bites, so seems like OpenAI implementation is just bad because o3 hallucination benchmark shouldn’t lead to such poor performance

Re: Introducing deep research

#404

Earlier quoted context omitted.

"Pretty neat, but definitely watch out for hallucinations." We'd never hire someone who just makes stuff up (or at least keep them employed for long). Why are we okay with calling "AI" tools like this anything other than curious research projects? Can't we just send LLMs back to the drawing board until they have some semblance of reliability?

> We'd never hire someone who just makes stuff up (or at least keep them employed for long). This is contrary to my experience.

Our president begs to differ! Or pretty much any elected official for that matter.

Re: Introducing deep research

#405
post #24

If I understood the graphs correctly, it only achieves 20% pass rate on their internal tests. So I have to wait 30min and pay a lot of money just to sift through walls of most likely incorrect text? Unless the possibility of hallucinations is negligible, this is just way too much content to review at once. The process probably needs to be a lot more iterative.

Here's an example of the type of question it is acheiving 20% on; The set of natural transformations between two functors F,G ⁣:C→DF,G:C→D can be expressed as the end Nat(F,G)≅∫AHomD(F(A),G(A)). Nat(F,G)≅∫A HomD (F(A),G(A)). Define set of natural cotransformations from FF to GG to be the coend CoNat(F,G)≅∫AHomD(F(A),G(A)). CoNat(F,G)≅∫AHomD (F(A),G(A)). Let: - F=B∙(Σ4)∗/F=B∙ (Σ4 )∗/ be the under ∞∞-category of the ne…

Do we actually know whether it got this specific example right? It got 20% on HLE, but I think a few questions are quite a bit easier.

Re: Introducing deep research

#406

I just gave it a whirl. Pretty neat, but definitely watch out for hallucinations. For instance, I asked it to compile a report on myself (vain, I know.) In this 500-word report (ok, I'm not that important, I guess), it made at least three errors. It stated that I had 47,000 reputation points on Stack Overflow -- quite a surprise to me, given my minimal activity on Stack Overflow over the years. I popped over to the l…

"Pretty neat, but definitely watch out for hallucinations." We'd never hire someone who just makes stuff up (or at least keep them employed for long). Why are we okay with calling "AI" tools like this anything other than curious research projects? Can't we just send LLMs back to the drawing board until they have some semblance of reliability?

> We'd never hire someone who just makes stuff up

We do all the time - of course we do, all the time.

Re: Introducing deep research

#407

I just gave it a whirl. Pretty neat, but definitely watch out for hallucinations. For instance, I asked it to compile a report on myself (vain, I know.) In this 500-word report (ok, I'm not that important, I guess), it made at least three errors. It stated that I had 47,000 reputation points on Stack Overflow -- quite a surprise to me, given my minimal activity on Stack Overflow over the years. I popped over to the l…

So, I still think this is a cool tool for search reasons, but otherwise the tendency to hallucinate makes it questionable as a researcher.

Hypothetically speaking, if the time you saved is now spent verifying the statements of your AI researcher, then did you really save any time at all?

If the answers aren't important enough to verify, then was it ever even important enough to actually research to begin with?

Re: Introducing deep research

#408

Earlier quoted context omitted.

Not sure if this was posted as humour, but I don't feel that way. In today's world, where I certainly would consider taking the blue pill, I'm having a blast with LLMs! It has helped me learn stuff incredibly faster. Especially I find them useful for filling the gaps of knowledge and exploring new topics in my own way and language, without needing to wait an answer from a human (that could also be wrong). Why does it…

Are you sure it's helped you learn? In the early days of ChatGPT where it seemed like this fun new thing, I used it to "learn" C. I don't remember anything it told me, and none of the answers it gave me were anything that I couldn't find elsewhere in different forms - heck I could have flipped open Kernighan & Ritchie to the right page and got the answer. I had a conversation with an AI/Bitcoin enthusiast recently. M…

I guess the real learning happens outside the AI, here in real life. Does the code run? Sure, it's on my local and not in production, but I would've never have the patience to get "that new thing working" without AI as assistant.

Does the food taste good? Oops, there's a bit too much vegetables here, they are never gonna fit in this pan of mine. Not a big deal, next time I'll be wiser.

AI is like a hypothesis machine. You're gonna have to figure out if the output is true. Few years ago, just testing any machine's "intelligence" was pretty quickly done and machine failed miserably. Now, the accuracy is astounishing in comparison.

> How many falsehoods influence you?

That is a great question. The answer is definitely not zero. I try to live by with a hacker mentality and I'm an engineer by trade. I read news and comments, which I'm not sure is good for me. But you also need some compassion towards oneself. It's not like ripping everything open will lead to salvation. I believe the truth does set you free, eventually. But all in one's time...

Anyway, AI is a tool like any other. Someone will hammer their fingers with it. I just don't understand the hate. It's not like we're drinking any AI koolaids here. It's just like it was 30 years ago (in my personal journey), you had a keyboard and a machine, you asked it things and got gibberish. Now the conversation with it just started to get interesting. Peace.

Re: Introducing deep research

#409

Earlier quoted context omitted.

"Pretty neat, but definitely watch out for hallucinations." We'd never hire someone who just makes stuff up (or at least keep them employed for long). Why are we okay with calling "AI" tools like this anything other than curious research projects? Can't we just send LLMs back to the drawing board until they have some semblance of reliability?

Why not just verify the output? It’s faster than generating the entire thing yourself. Why do you need perfection in a productivity tool?

At that point why not just... I dunno, do the research yourself?

Re: Introducing deep research

#410
post #333

Gemini has had this for a month or two, also named "Deep Research" https://blog.google/products/gemini/google-gemini-deep-resea... Meta question: what's with all of the naming overlap in the AI world? Triton (Nvidia, OpenAI) and Gro{k,q} (X.ai, groq, OpenAI) all come to mind

> what's with all of the naming overlap in the AI world? Triton (Nvidia, OpenAI) and Gro{k,q} (X.ai, groq, OpenAI) all come to mind They seem to be ok with outsourcing any and all creativity to a language model, so it’s not surprising that they can’t come up with unique names themselves.

lol
Post reply on HN