Live data from Hacker News

We can do better than DuckDuckGo

drewdevault.com

81–90 of 383 posts

Re: We can do better than DuckDuckGo

#81

Am I the only person who just doesn't have problems with DDG search results? What am I doing wrong (or right), here? I put a thing in and find it. I just don't use Google any more. Genuinely curious why it's working for me and such garbage for everyone else.

I'm mostly getting Norwegian results, when searching for Danish subjects from a Danish IP address. It also seems it just hasn't indexed as many websites as Google.

Re: We can do better than DuckDuckGo

#82
post #66

Earlier quoted context omitted.

In theory , this is the kind of thing that the GPL v3 was trying to address: roughly speaking, if you host & run a service that is derived from GPL-v3'd software, you are obliged to publish your modifications. But, I agree with you - and I don't think the author had really thought through what they were demanding, they made no mention of licensing other than singing happy praises of FOSS as if that would magically me…

> In theory, this is the kind of thing that the GPL v3 was trying to address: roughly speaking, if you host & run a service that is derived from GPL-v3'd software, you are obliged to publish your modifications. You mean AGPL https://en.m.wikipedia.org/wiki/Affero_General_Public_Licens...

You're right... I'm misremembering the GPL, wikipedia says that it was only 'Early drafts of GPLv3 also let licensors add an Affero-like requirement that would have plugged the ASP loophole in the GPL' - I hadn't realised it never made it into the final version.

Re: We can do better than DuckDuckGo

#83
post #51
post #42

Earlier quoted context omitted.

How would anyone block a crawler? A crawler is just a headless browser.

robots.txt https://www.robotstxt.org/ https://en.wikipedia.org/wiki/Robots_exclusion_standard

Note that robots.txt is a hint to well-behaved crawlers, not blocking them in any regard.

You can block crawlers if you can identify them, but reliably identifying them is hard.

Re: We can do better than DuckDuckGo

#85
Yes, we can do better than DDG. But if you are expecting to fund a real search engine with a few hundred thousand dollars you are insane. It will take a ton of development and a ton of hardware to create an index that isn't a pile of garbage. This isn't 2000 anymore. You need to index >100 billion pages and you need it updated and you need great crawling and parsing and you need great algorithms and probably an entirely proprietary engine and you need to CONSTANTLY refine all the above until it isn't garbage. Maybe you could muster something passable for $1B over 5 years with a strong core team that attracts great talent. If Apple actually does this, as they are rumored to, I bet they dump $10b into it just for the initial version.

Re: We can do better than DuckDuckGo

#86
post #80

Privacy or not I'm starting to find things on ddg that google has been filtering. I found out through comments on hn that 8chan was backup under a new name: 8kun Typing it into google I get articles about it but no link in the results. In duckduckgo first link. Made me think what else am I missing?

Googles filtering is so weird. I have a bad habit of buying old hardware without checking if the documentation had been made available on the web.

Recently I found myself desperate for any information on a price of hardware i had gotten. I was swapping out all sorts of queries woth different keywords hoping to find a manual. I was able to find some marketing material which was helpful, albeit barely. Eventually I had exhausted the search results for most pf my queries, gave up and assumed that it was simply lost to time and I was out of luck.

Eventually I went back to the sales paper I found. Going to the site it was hosted on, a Lithuanian reseller. I translated the page, eventually finding a direct link to a user manual on the exact same page as the sales paper I had found. The document was in English, contained important words from my queries (such as the product name, company, "user manual" etc. The document was at the same path as the sales paper too. I hace no idea why Google found the sales paper but not the manual.

Unfortunately the manual still wasn't what I was looking for exactly but it was a hell of a lot better than what I could get from Google's results.

Re: We can do better than DuckDuckGo

#87
post #21

> they’ve demonstrated gross incompetence in privacy Not sure I buy the example that is given here. 1. It's an issue in their browser app, not their search service. 2. It's not completely indefensible: it allows fetching favicons (potentially) much faster, since they're cached, and they promise that the favicon service is 100% anonymous anyway. 3. They responded to user feedback and switched to fetching favicons loca…

Agreed. I think the key point here is that the web is a radically different place than it was in 1998 (when Google launched and established the search engine paradigm as we know it). Back then the quality-to-spam ratio was probably much higher, the overall size of the web was certainly much smaller (making scraping the entire thing more tractable), and there were many more self-hosted sources rather than platforms (meaning it was more necessary to rely on inter-linking, and "authoritative domains" weren't as much of a thing). The naive scraping approach was both more crucial and more effective. And in the decades since, it's been a constant war of attrition to keep that model working under more and more adversarial conditions.

So I think that stepping back and re-thinking what a search engine fundamentally is, is a great starting point for disruption.

Additionally, something the OP didn't mention is that ML technologies have progressed dramatically since 1998, and that much of that progress has been done in the open. I can't imagine that not being a force-multiplier for any upstart in this domain.

Re: We can do better than DuckDuckGo

#88

Am I the only person who just doesn't have problems with DDG search results? What am I doing wrong (or right), here? I put a thing in and find it. I just don't use Google any more. Genuinely curious why it's working for me and such garbage for everyone else.

Does Google search results work for you? If yes, then I'd say the reason is you don't see or agree with how bad results are today (as others have posted extensively about). I for one find DDG as the search engine that returns the worst results. Qwant is a better Bing-using engine IMO but it is still bad.

Re: We can do better than DuckDuckGo

#89
post #20

Earlier quoted context omitted.

I don’t even care about the privacy. (Well, I do, but in this context I have no reasonable way to ensure it) What I do care about is trust-building and monopolistic practices. That, to me, is a great reason to use DDG instead of Google or even Bing.

I also prefer DDG's user interface over Google's. And DDG's !bang search shortcuts. DDG has been my default search engine for years and its results are good enough for me 95% of the time. I only need to use Google as a fallback when searching for niche technical information or "needles in haystacks".

Even then, the Google results are usually terrible. I haven't used Google as a fallback in about a year because every time I tried it they couldn't find what I was looking for either. Or they did something atrocious like changing my search terms for me.

Re: We can do better than DuckDuckGo

#90
post #34
post #21

> they’ve demonstrated gross incompetence in privacy Not sure I buy the example that is given here. 1. It's an issue in their browser app, not their search service. 2. It's not completely indefensible: it allows fetching favicons (potentially) much faster, since they're cached, and they promise that the favicon service is 100% anonymous anyway. 3. They responded to user feedback and switched to fetching favicons loca…

Maybe instead of hard-coding these preferences in the search engine, or having it try to guess for you based on your search history, you can opt-in to download and apply such lists of ranking modifiers to your user profile. Those lists would be maintained by 3rd parties and users, just like eg. adblock blacklists and whitelists. For example, Python devs might maintain a list of search terms and associated urls that g…

This is a great idea. It's like a modern reboot of the old concept of curated "link lists", maintained by everyone from bloggers to Yahoo. Doing it at a meta level for search-engine domains is a really cool thought.
Post reply on HN