Live data from Hacker News

Introducing deep research

openai.com

331–340 of 445 posts

Re: Introducing deep research

#333

Gemini has had this for a month or two, also named "Deep Research" https://blog.google/products/gemini/google-gemini-deep-resea... Meta question: what's with all of the naming overlap in the AI world? Triton (Nvidia, OpenAI) and Gro{k,q} (X.ai, groq, OpenAI) all come to mind

> what's with all of the naming overlap in the AI world? Triton (Nvidia, OpenAI) and Gro{k,q} (X.ai, groq, OpenAI) all come to mind

They seem to be ok with outsourcing any and all creativity to a language model, so it’s not surprising that they can’t come up with unique names themselves.

Re: Introducing deep research

#334

Gemini has had this for a month or two, also named "Deep Research" https://blog.google/products/gemini/google-gemini-deep-resea... Meta question: what's with all of the naming overlap in the AI world? Triton (Nvidia, OpenAI) and Gro{k,q} (X.ai, groq, OpenAI) all come to mind

I am afraid Gemini's version is not really very "deep" - it surfaces a lot of information, but on a quite superficial level. OAIs version seems to make that one step forward to proper depth.

We found in our experience it is pretty hard to force LLM to do something in proper depth, and OAI's deep research definitely feels like one of the first examples from big labs on how this can be done. What we typically see is that it is not even the "agent" part that is hard to do, but how to force model to not "forget" to go deep...

Re: Introducing deep research

#338

I just gave it a whirl. Pretty neat, but definitely watch out for hallucinations. For instance, I asked it to compile a report on myself (vain, I know.) In this 500-word report (ok, I'm not that important, I guess), it made at least three errors. It stated that I had 47,000 reputation points on Stack Overflow -- quite a surprise to me, given my minimal activity on Stack Overflow over the years. I popped over to the l…

What's faster? Writing a 500 word report "from scratch" by researching the topic yourself, vs. having AI write it then having to fact check every answer and correct each piece manually?

This is why I don't use AI for anything that requires a "correct" answer. I use it to re-write paragraphs or sentences to improve readability etc, but I stop short of trusting any piece of info that comes out from AI.

Re: Introducing deep research

#339

Earlier quoted context omitted.

It contributes little to discuss a hypothetical future. Maybe we'll have fusion energy, delivery drones, everyone using VR, etc. Maybe we will go into a deep recession due to trade wars, or maybe not. The meaningful discussion is about how they perform NOW and the edge cases that have persisted since GPT-2 which no one has yet found a good solution for.

We already have delivery drones though. I disagree though, it is useful as this problem has been whittled down and I think there is expectation that there will be continued effort. Its of course worth discussing but I find that for my workflows, I rarely encounter issues with hallucinations, they certainly exist but its gotten to a point that I don't have major issue with it.

At best, a proof of concept of experimental delivery drones exist, but only for small, lightweight items, and only in a few places, only in the right weather, and only if you place a target on your driveway and are there to receive the item in person, and all at the cost of a very high noise level. That's not exactly a real service.

Re: Introducing deep research

#340

I just gave it a whirl. Pretty neat, but definitely watch out for hallucinations. For instance, I asked it to compile a report on myself (vain, I know.) In this 500-word report (ok, I'm not that important, I guess), it made at least three errors. It stated that I had 47,000 reputation points on Stack Overflow -- quite a surprise to me, given my minimal activity on Stack Overflow over the years. I popped over to the l…

> Then it cited an answer I gave on Stack Overflow [...] using this as evidence for my technical expertise. Turns out that about 15 years ago, I _asked_ a question on this topic, but the answer was submitted by someone else

Artificial dementia...

Some parties are releasing products much earlier than the ability to ship well working products (I am not sure that their legal cover will be so solid), but database aided outputs should and could become a strong limit to that phenomenon of remembering badly. Very linearly, like humans: get an idea, then compare it to the data - it is due diligence and part of the verification process in reasoning. It is as if some moves outside linear pure product progress reasoning are swaying the RnD towards directions outside the primary concerns. It's a form of procrastination.

Post reply on HN