Live data from Hacker News

Open Deep Research

github.com

61–70 of 83 posts

Re: Open Deep Research

#61
post #11

https://techcrunch.com/2025/02/04/hugging-face-researchers-a... > On GAIA, a benchmark for general AI assistants, Open Deep Research achieves a score of 54%. That’s compared with OpenAI deep research’s score of 67.36%..Worth noting is that there are a number of OpenAI deep research “reproductions” on the web, some of which rely on open models and tooling. The crucial component they — and Open Deep Research — lack is…

theres always a lot of openTHING clones of THING after THING is announced. they all usually (not always[1]!) disappoint/dont get traction. i think the causes are 1. running things in production/self hosting is more annoying than just paying like 20-200/month 2. openTHING makers often overhype their superficial repros ("I cloned Perplexity in a weekend! haha! these VCs are clowns!") and trivializing the last mile, mos…

As for them not getting traction - many projects just don’t advertise which tech they use.

Also, having an open source alternative - even if slightly worse - gives projects an alternative if OpenAI decides to cut them off or sth.

What I tell startups is to start off with whatever OpenAI/Anthropic have to offer, if they can, and consider switching into Open Source only if it allows for something they can’t do (often in gen graphics), or when they reach product market fit and they can finetune/train a smaller model that handles their specific case better and cheaper.

Re: Open Deep Research

#62
post #17
post #12

Earlier quoted context omitted.

The rise of captchas on regular content, no longer just for posting content, could ruin this. Cloudflare and other companies have set things up to go through a few hand selected scrapers and only they will be able to offer AI browsing and research services.

I think the opposite problem is going to occur with captchas for whatever it's worth: LLMs are going to obsolete them. It's an arms race where the defender has a huge constraint the attacker doesn't (pissing off real users); in that way, it's kind of like the opposite dynamics that password hashes exploit.

Cloudflare is more than captchas, it's centralized monitoring of them too: what do you think happens when your research assistant solves 50 captchas in 5 min from your home IP? It has to slow down to human research speeds.

Re: Open Deep Research

#63

I signed up for Gemini Advanced to get access to Deep Research on Feb 1st and felt on top of the world. So cool, so advanced. Then OpenAI announced theirs on the 2nd: https://openai.com/index/introducing-deep-research/ Ethan Mollick called Google's undergraduate level and OpenAI's graduate level on the 3rd: https://www.oneusefulthing.org/p/the-end-of-search-the-begin... And now this. I can't stop thinking about The O…

My X feed is 90% AI Influencers like this saying "I'm blown away! Check this out.." Is this a new cottage industry? Are they making money?

Yea x and YouTube. Also the YouTube meta specifically seems to have moved to these sensationalist over exaggeration and over simplification. Really would be nice to have a sensibility filter somehow. It’s gonna come with ai.

Re: Open Deep Research

#64
post #17

Earlier quoted context omitted.

I think the opposite problem is going to occur with captchas for whatever it's worth: LLMs are going to obsolete them. It's an arms race where the defender has a huge constraint the attacker doesn't (pissing off real users); in that way, it's kind of like the opposite dynamics that password hashes exploit.

I’m not sure about that. There’s a lot of runway left for obstacles that are easy for humans and hard/impossible for AI, such as direct manipulation puzzles. (AI models have latency that would be impossible to mask.) On the other hand, a11y needs do limit what can be lawfully deployed…

There’s a lot of runway left for obstacles that are easy for humans and hard/impossible for AI, such as direct manipulation puzzles.

That's irrelevant. Humans totally hate CAPTCHAs and they are an accessibility and cultural nightmare. Just forget about them. Forget about making better ones, forget about what AI can and can't do. We moved on from CAPTCHAs for all those reasons. Everyone else needs to.

Re: Open Deep Research

#65

Earlier quoted context omitted.

You're being slippery with language. The open source version doesn't get 26% on Humanity's Last Exam. It's not "already out".

> Humanity's Last Exam Good grief, AI researchers need to learn basic humility if they want to market their tech successfully. Unless they're toppling our states and liberating us from capital (extremely difficult to imagine) I have a difficult time imagining any value or threat AI could provide that necessitates this level of drama.

Constantly warning about how your product is or will eventually be so powerful that it will become a danger to humanity or society is the successful marketing.

And for the big AI companies, like OpenAI, it has the very beneficial side-effect of establishing the narrative that lets them influence politics into regulating their potential competitors out of the market. Because they are, of course, the only reasonable and responsible builders of self-described doomsday devices.

Re: Open Deep Research

#66
post #61
post #11

Earlier quoted context omitted.

theres always a lot of openTHING clones of THING after THING is announced. they all usually (not always[1]!) disappoint/dont get traction. i think the causes are 1. running things in production/self hosting is more annoying than just paying like 20-200/month 2. openTHING makers often overhype their superficial repros ("I cloned Perplexity in a weekend! haha! these VCs are clowns!") and trivializing the last mile, mos…

As for them not getting traction - many projects just don’t advertise which tech they use. Also, having an open source alternative - even if slightly worse - gives projects an alternative if OpenAI decides to cut them off or sth. What I tell startups is to start off with whatever OpenAI/Anthropic have to offer, if they can, and consider switching into Open Source only if it allows for something they can’t do (often i…

> if it allows for something they can’t do (often in gen graphics)

What makes gen graphics stand out?

Re: Open Deep Research

#67
post #66
post #61

Earlier quoted context omitted.

As for them not getting traction - many projects just don’t advertise which tech they use. Also, having an open source alternative - even if slightly worse - gives projects an alternative if OpenAI decides to cut them off or sth. What I tell startups is to start off with whatever OpenAI/Anthropic have to offer, if they can, and consider switching into Open Source only if it allows for something they can’t do (often i…

> if it allows for something they can’t do (often in gen graphics) What makes gen graphics stand out?

I’d venture to say it’s the range:

What style? Is it a texture? If so, I’ll need a model that can generate a large tiled image. Is it a logo? What kind? Is it badged, vintage, corporate, etc

Re: Open Deep Research

#68
post #15
post #11

Earlier quoted context omitted.

theres always a lot of openTHING clones of THING after THING is announced. they all usually (not always[1]!) disappoint/dont get traction. i think the causes are 1. running things in production/self hosting is more annoying than just paying like 20-200/month 2. openTHING makers often overhype their superficial repros ("I cloned Perplexity in a weekend! haha! these VCs are clowns!") and trivializing the last mile, mos…

Maybe these open projects start to get more attention when we have a distribution system/App Store for AI projects. I know YC is looking to fund this https://www.ycombinator.com/rfs

[deleted]

Re: Open Deep Research

#69
post #11

Earlier quoted context omitted.

theres always a lot of openTHING clones of THING after THING is announced. they all usually (not always[1]!) disappoint/dont get traction. i think the causes are 1. running things in production/self hosting is more annoying than just paying like 20-200/month 2. openTHING makers often overhype their superficial repros ("I cloned Perplexity in a weekend! haha! these VCs are clowns!") and trivializing the last mile, mos…

> running things in production/self hosting is more annoying than just paying like 20-200/month This is an important point. As people rely more on AI/LLM tools, reliability will become even more critical. In the last two weeks, I've heavily used Claude and DeepSeek Chat. ChatGPT is much more reliable compared to both. Claude struggles with long-context chats and often shifts to concise responses. DeepSeek often has i…

> In the last two weeks, I've heavily used Claude and DeepSeek Chat. ChatGPT is much more reliable compared to both

Which reliability problems did you face? I heard about connection issues due to too much traffic with Deepseek, but those would go away if you self-host the model.

Re: Open Deep Research

#70

Earlier quoted context omitted.

I’m not sure about that. There’s a lot of runway left for obstacles that are easy for humans and hard/impossible for AI, such as direct manipulation puzzles. (AI models have latency that would be impossible to mask.) On the other hand, a11y needs do limit what can be lawfully deployed…

There’s a lot of runway left for obstacles that are easy for humans and hard/impossible for AI, such as direct manipulation puzzles. That's irrelevant. Humans totally hate CAPTCHAs and they are an accessibility and cultural nightmare. Just forget about them. Forget about making better ones, forget about what AI can and can't do. We moved on from CAPTCHAs for all those reasons. Everyone else needs to.

Agreed. When I open a link and get a Cloudflare CAPTCHA I just close the tab.
Post reply on HN