Live data from Hacker News

Introducing deep research

openai.com

191–200 of 445 posts

Re: Introducing deep research

#191
post #72

Earlier quoted context omitted.

Anyone selling anything would want to remain crawlable if people use this to research something that could lead to a purchase.

Not necessarily. Southwest airlines doesnt allow itself on price comparison sites or Google Flights. Amazon listings are blocked from google shopping and other price comparison sites.

Your point is completely valid, but... Southwest now has an arrangement with Google Flights to allow their listings there.

Re: Introducing deep research

#193

Gemini has had this for a month or two, also named "Deep Research" https://blog.google/products/gemini/google-gemini-deep-resea... Meta question: what's with all of the naming overlap in the AI world? Triton (Nvidia, OpenAI) and Gro{k,q} (X.ai, groq, OpenAI) all come to mind

I've always thought the Triton situation was intentional since the name isn't generic and because the companies are stepping on each others toes here (Nvidia's Triton simplifying owning your inference; OpenAI's Triton eroding the need for familiarity with CUDA). I couldn't figure out who publicly used the name first though.

Re: Introducing deep research

#194
"Deep research" is now somehow synonymous to searching online for stats and pulling stuff from Statista? And when I want to make changes to that report, do I have to tweak my prompt and get an entirely different document?

Not sure if I'm too tired and can't see it but the lack of images/examples of the resulting report in this announcement doesn't inspire a lot of confidence just yet.

Re: Introducing deep research

#195

This smells like when Google released Gemini to have a product in the space.

I agree that OpenAI is trying to stay relevant by announcing a lot of have baked products with little to no availability.

> when Google released Gemini to have a product in the space.

Bard preceded Gemini.

Re: Introducing deep research

#196
Can it compile and run (non-Python) code as part of its tool use? Compile-run steps always seemed like they would be a huge value add during reasoning loops - it feels very silly to get output from ChatGPT, try to run it in terminal, get an error and paste the error to have ChatGPT immediately fix it. Surely it should be able to run code during the reasoning loop itself?

Re: Introducing deep research

#197

Earlier quoted context omitted.

It's a direction in a vast landscape, not a feature of itself - being better at different tasks, like search generally, and research in conjunction with reasoning, gets the model closer to AGI. An AGI will be able to do these tasks - so the point of the research is to have more Venn diagrams of capabilities like these to help narrow down the view on things that might actually be fundamental mechanisms involved in AGI…

> hings like feeling pain and pleasure can machine feel? without that there is no AGI according to definition above. and the second question: are animals "GI"? they don't have language and don't solve math problems, never heard of np-complete.

Are we not machines anyway ? Ofc a machine can feel, just need to have priorities that are aligned to itself, and use strong feedback when that self is either in danger or on the right path to preservation...

Feelings are nothing very special you know...

Re: Introducing deep research

#198
Each release from openAI gives me less hope for them and this whole AI boom. They should be leading the charge of highlighting how the current generation of LLMs fail, not churning out half-baked overhyped products.

Yes, they can do some cool tricks, and tool calling is fun. No one should trust the output of these models, though. The hallucinations are bad, and my experience with the "reasoning" models is that as soon as they fuck up (they always do) they go off the rails worse than the base LLMs.

Re: Introducing deep research

#199
post #24

If I understood the graphs correctly, it only achieves 20% pass rate on their internal tests. So I have to wait 30min and pay a lot of money just to sift through walls of most likely incorrect text? Unless the possibility of hallucinations is negligible, this is just way too much content to review at once. The process probably needs to be a lot more iterative.

On questions even specialists in that field can’t answer correctly.

Re: Introducing deep research

#200
post #128

Earlier quoted context omitted.

Its number for one of the benchmark has: **with browsing + python tools Maybe we have different definitions of scaling?

I would consider unsupervised tool usage an achievement

But it’s not simply scaling. Who is moving the goalposts exactly?
Post reply on HN