To anyone who's tried it: how does it handle captchas? I can't imagine that OpenAI's IP addresses are anyone's favorites for unfettered access to web properties these days.
Introducing deep research
131–140 of 445 posts
Re: Introducing deep research
#132Not sure if people picked up on it, but this is being powered by the unreleased o3 model. Which might explain why it leaps ahead in benchmarks considerably and aligns with the claims o3 is too expensive to release publicly. Seems to be quite an impressive model and the leading out of Google, DeepSeek and Perplexity.
It was expensive as they wanted to charge more for it but deepseek has forced their hand
Re: Introducing deep research
#133Reasoning, problem solving, research validation - at the fundamental outset it is all refinement thinking.
Research is one of those areas where I remain skeptical it is that important because the only valid proof is in the execution outcome, not the compiled answer.
For instance you can research all you want about the best vacuum on the internet but until you try it out yourself you are going to be caught in between marketing, fake reviews, influencers, etc. maybe the science fields are shielded from this (by being boring) but imagine medical pharmas realizing that they can get whatever paper to say whatever by flooding the internet with their curated blog articles containing advanced medical “research findings”. At some point you cannot trust the internet at all and I imagine that might be soon.
I worry especially with the rapidly changing landscape of the amount of generated text in the internet that research will lose a lot of value due to massive amounts of information garbage.
It will be a thing we used to do when the internet was still “real”.
Re: Introducing deep research
#134So much cynicism and hate in these comments, especially as we are likely witnessing AGI come to life. Its still early, but it might be coming. Where is the excitement? This is an interesting time to be alive. HN has a huge cultural problem that makes this website almost irrelevant. All the interesting takes have moved to X/twitter
More seriously, it’s unclear why one should be excited by the prospect of AGI, especially when instrumentalized by corporations and authoritarian governments.
Re: Introducing deep research
#135Not sure if people picked up on it, but this is being powered by the unreleased o3 model. Which might explain why it leaps ahead in benchmarks considerably and aligns with the claims o3 is too expensive to release publicly. Seems to be quite an impressive model and the leading out of Google, DeepSeek and Perplexity.
Re: Introducing deep research
#136Surprised more comments aren't mentioning deepseek has this feature (for free) already. Assuming this is why OpenAI scrambled to release it. The examples they have on the page work well on chat.deepseek.com with r1 and search options both enabled. Do I blindly trust the accuracy of either though? Absolutely not. I'm pretty concerned about these models falling into gaming SEO and finding inaccurate facts and presentin…
Not really accurate. The "Search" functionality you're describing in DeepSeek is comparable to OpenAI's existing "Search GPT." OpenAI's recent announcement refers to a more advanced capability, similar to Gemini's existing "deep research" feature. DeepSeek's current offerings are significantly more limited in scope.
AFAIK OpenAI's current offering uses 4o, and it does a web search and then pipes it into 4o. I'm guessing adding CoT + other R1/o3 like stuff is one of the key effective differences. But time will tell how different it is. Maybe it's a dramatic improvement.
Re: Introducing deep research
#137Earlier quoted context omitted.
Either you care about being correct or you don't. If you don't care then it doesn't matter whether you made it up or the AI did. If you care then you'll fact check before publishing. I don't see why this changes.
I think a lot about how differentiating facts and quality content is like differentiating signal from noise in electronics. The signal to noise ratio on many online platforms was already quite low. Tools like this will absolutely add more noise, and arguably the nature of the tools themselves make it harder to separate the noise. I think this is a real problem for these AI tools. If you can’t separate the signal from…
Re: Introducing deep research
#138If I understood the graphs correctly, it only achieves 20% pass rate on their internal tests. So I have to wait 30min and pay a lot of money just to sift through walls of most likely incorrect text? Unless the possibility of hallucinations is negligible, this is just way too much content to review at once. The process probably needs to be a lot more iterative.
26.6% on humanity's last exam is actually impressive. pass rate really only matters in context of the difficulty of the tasks
Re: Introducing deep research
#139Surprised more comments aren't mentioning deepseek has this feature (for free) already. Assuming this is why OpenAI scrambled to release it. The examples they have on the page work well on chat.deepseek.com with r1 and search options both enabled. Do I blindly trust the accuracy of either though? Absolutely not. I'm pretty concerned about these models falling into gaming SEO and finding inaccurate facts and presentin…
Not really accurate. The "Search" functionality you're describing in DeepSeek is comparable to OpenAI's existing "Search GPT." OpenAI's recent announcement refers to a more advanced capability, similar to Gemini's existing "deep research" feature. DeepSeek's current offerings are significantly more limited in scope.
Re: Introducing deep research
#140I'm sorry but what the fuck is this product pitch? Anyone who's done any kind of substantial document research knows that it's a NIGHTMARE of chasing loose ends & citogenesis. Trusting an LLM to critically evaluate every source and to be deeply suspect of any unproven claim is a ridiculous thing to do. These are not hard reasoning systems, they are probabilistic language models.