Live data from Hacker News

Open Deep Research

github.com

51–60 of 83 posts

Re: Open Deep Research

#51
post #42

Earlier quoted context omitted.

I’m not sure about that. There’s a lot of runway left for obstacles that are easy for humans and hard/impossible for AI, such as direct manipulation puzzles. (AI models have latency that would be impossible to mask.) On the other hand, a11y needs do limit what can be lawfully deployed…

>AI models have latency So do humans, or can my friend with cerebral palsy not use the internet any longer?

Totally different type of latency. A person with a motor disability dragging a puzzle piece with their finger will look very different from an AI model being called frame by frame.

Re: Open Deep Research

#52
post #43
post #31

Hi all! Aymeric (m-ric) here, maintainer of smolagents and part of the team who built this. Happy to see this interesting people here! Few points: - open Deep Research is not a production app, but it could easily be productionized (would need to be faster + good UX). - As the GAIA score of 55% (not 54%, that would be lame) says, it's not far from the Deep Research score of 67%. It's also not there yet: I think the ma…

I think using vision models for browsing is the wrong approach. It is the same as using OCR for scanning PDFs. The underlying text is already in digital form. So it would make more sense to establish a standard similar to meta-tags that enable the agentic web.

If you're working from the markup rather than the appearance of the page, you're probably increasing the incentives for metacrap, "invisible text spam" and similar tactics.

Re: Open Deep Research

#53
post #32

Earlier quoted context omitted.

> Nothing that Perplexity + DeepSeek-R1 can already do Any public comparisons of OAI Deep Research report quality with Perplexity + DeepSeek-R1, on the same query? How do cost and query limits compare?

I've been using GenSpark.ai for the past month to do research (its agents usually does ~20 minutes, but I've seen it go up to almost 2 hours on a task) - it uses a Mixture of Agents approach using GPT-4o, Claude 3.5 Sonnet and Gemini 1.5 Pro and searches for hundreds of sources. I reran some of these searches and I've so far found OpenAI Deep Research to be superior for technical tasks. Here's one example: https://ch…

> still mystified if switching between the different base model matters, besides o1 pro always seeming to fail to execute the Deep Research tool.

You mean when it says it's going to research and get back to you and then ... just doesn't?

Re: Open Deep Research

#54
post #32

Earlier quoted context omitted.

I've been using GenSpark.ai for the past month to do research (its agents usually does ~20 minutes, but I've seen it go up to almost 2 hours on a task) - it uses a Mixture of Agents approach using GPT-4o, Claude 3.5 Sonnet and Gemini 1.5 Pro and searches for hundreds of sources. I reran some of these searches and I've so far found OpenAI Deep Research to be superior for technical tasks. Here's one example: https://ch…

> still mystified if switching between the different base model matters, besides o1 pro always seeming to fail to execute the Deep Research tool. You mean when it says it's going to research and get back to you and then ... just doesn't?

Yeah, it seems to not be able to execute the tool calling properly. Maybe it's a bad interaction w/ it's own async calling ability or something else (eg, how search and code interpreter can't seem to run at the same time for 4o)

Re: Open Deep Research

#55

I signed up for Gemini Advanced to get access to Deep Research on Feb 1st and felt on top of the world. So cool, so advanced. Then OpenAI announced theirs on the 2nd: https://openai.com/index/introducing-deep-research/ Ethan Mollick called Google's undergraduate level and OpenAI's graduate level on the 3rd: https://www.oneusefulthing.org/p/the-end-of-search-the-begin... And now this. I can't stop thinking about The O…

My X feed is 90% AI Influencers like this saying "I'm blown away! Check this out.."

Is this a new cottage industry? Are they making money?

Re: Open Deep Research

#56

I signed up for Gemini Advanced to get access to Deep Research on Feb 1st and felt on top of the world. So cool, so advanced. Then OpenAI announced theirs on the 2nd: https://openai.com/index/introducing-deep-research/ Ethan Mollick called Google's undergraduate level and OpenAI's graduate level on the 3rd: https://www.oneusefulthing.org/p/the-end-of-search-the-begin... And now this. I can't stop thinking about The O…

My X feed is 90% AI Influencers like this saying "I'm blown away! Check this out.." Is this a new cottage industry? Are they making money?

I’ve heard it’s all the crypto influencers jumping into the AI hype

Re: Open Deep Research

#57
post #43
post #31

Hi all! Aymeric (m-ric) here, maintainer of smolagents and part of the team who built this. Happy to see this interesting people here! Few points: - open Deep Research is not a production app, but it could easily be productionized (would need to be faster + good UX). - As the GAIA score of 55% (not 54%, that would be lame) says, it's not far from the Deep Research score of 67%. It's also not there yet: I think the ma…

I think using vision models for browsing is the wrong approach. It is the same as using OCR for scanning PDFs. The underlying text is already in digital form. So it would make more sense to establish a standard similar to meta-tags that enable the agentic web.

PDFs are more akin to SVG than to a Word document, and the text is often very far from “available”. OCR can be the only way to reconstruct the document as it appears on screen.

Re: Open Deep Research

#58
post #13

Of course. The first of many open source versions of 'Deep Research' projects are now appearing as predicted [0] but in less than a month. Faster than expected. Open source is already at the finish line. [0] https://news.ycombinator.com/item?id=42913379

> The first of many open source versions of 'Deep Research' projects are now appearing as predicted [0] but in less than a month. Faster than expected.

Well, in that particular case the open source version was actually here first three month ago[1].

[1]: https://www.reddit.com/r/LocalLLaMA/comments/1gvlzug/i_creat...

Re: Open Deep Research

#59
Somehow "deep" has acquired the connotation of "bullshit". It's a shame—I'm very bullish on LLM technology, but capital seems determined to diminish every angle of potential by blatantly lying about its capabilities.

Granted, this isn't an entrepreneurial venture, so maybe it has some value. but this still stinks to high heaven of saving money rather than producing value. One day we'll see AI produce stuff of value that humans can't already do better (aside from playing board games), but that day is still a long way off.

Re: Open Deep Research

#60

So basically, Altman announced Deep Research less than a month ago and open-source alternatives are already out? Investors are not going to be happy unless OpenAI outperforms them all by an order of magnitude

You're being slippery with language. The open source version doesn't get 26% on Humanity's Last Exam. It's not "already out".

> Humanity's Last Exam

Good grief, AI researchers need to learn basic humility if they want to market their tech successfully. Unless they're toppling our states and liberating us from capital (extremely difficult to imagine) I have a difficult time imagining any value or threat AI could provide that necessitates this level of drama.

Post reply on HN