Live data from Hacker News

We can do better than DuckDuckGo

drewdevault.com

171–180 of 383 posts

Re: We can do better than DuckDuckGo

#171
post #21

> they’ve demonstrated gross incompetence in privacy Not sure I buy the example that is given here. 1. It's an issue in their browser app, not their search service. 2. It's not completely indefensible: it allows fetching favicons (potentially) much faster, since they're cached, and they promise that the favicon service is 100% anonymous anyway. 3. They responded to user feedback and switched to fetching favicons loca…

Some kind of logic like "For Python programming queries, docs.python.org and then StackExchange are the tier 1 sources" seems to be the kind of hard-coded information that would vastly improve my experience trying to look things up on DuckDuckGo.

The problem with this strategy is always going to be that different users will regard different sources as most desirable.

For example, it's enormously frustrating that searching for almost anything Python-related on DDG seems to return lots of random blog posts but hardly ever shows the official Python docs near the top. I don't personally think the official Python docs are ideally presented, but they're almost certainly more useful to me at that time than some random blog that happens to mention an API call I'm looking up.

On the other hand, I would gladly have an option in a search engine to hide the entire Stack Exchange network by default. The signal/noise ratio has been so bad for a long time that I would prefer to remove them from my search experience entirely rather than prioritise them. YMMV, of course. (Which is my point.)

Re: We can do better than DuckDuckGo

#173
post #34
post #21

> they’ve demonstrated gross incompetence in privacy Not sure I buy the example that is given here. 1. It's an issue in their browser app, not their search service. 2. It's not completely indefensible: it allows fetching favicons (potentially) much faster, since they're cached, and they promise that the favicon service is 100% anonymous anyway. 3. They responded to user feedback and switched to fetching favicons loca…

Maybe instead of hard-coding these preferences in the search engine, or having it try to guess for you based on your search history, you can opt-in to download and apply such lists of ranking modifiers to your user profile. Those lists would be maintained by 3rd parties and users, just like eg. adblock blacklists and whitelists. For example, Python devs might maintain a list of search terms and associated urls that g…

I have had a similar idea, what you're proposing is essentially a ranking/filtering customisation. The internet is a big scene, and on this scene we have companies and their products, political parties, ad agencies and regular users. Everyone is fighting for attention, clicks. Google has control over a ranking and filtering system that covers most searches on the internet. FB and Twitter hold another ranking/filtering sweet spot for social networks.

The problem is that we have no say in ranking and filtering. I think it should be customisable both on a personal and community level. We need a way to filter out the crap and surface the good parts on all these sites. I am sure Google wouldn't like to lose control of ranking and filtering, but we can't trust a single company with such an essential function of our society, and we can't force a single editorial view on everyone.

As we have many newspapers, each with its own editorial views, we need multiple search engine curators as well.

Re: We can do better than DuckDuckGo

#175
I agree DDG isn't perfect or great but it's _good_ 80% of the time.

I always start with DDG and revert back to Google if it doesn't help, or I feel "there's got to be a better way".

That said, talk is cheap, show us your engine.

Re: We can do better than DuckDuckGo

#176
What about a curated search engine? Allow anyone to curate the results and define a custom list of allowed URLs. Then, others can use their list.

For example, I decide Google is terrible when I'm searching for product reviews, and all I get are results to Amazon referral websites and spam blogs that never owned the products to begin with. So, I find 200 sites or forums that actually have quality reviews and I create a whitelist of those URLs, and I name it "John Doe's product reviews list".

Other people visit the search engine and they can see my list, favorite it, and apply it to their results. I maintain the list, so they continue to get updates as it's refined.

The idea is you visit the search engine, type your query, then select from a drop down one of your favorite curated lists to apply. Maybe you like to use "Mike's favorite free stock photo websites" when searching for free photos for your projects. Maybe you like to apply "Jane's vegan friendly results" when searching recipes or face creams. Maybe you want to buy local, so you use the "Handmade in X" list when searching for your next belt. Maybe you use a list that only shows results from forums, or another for tracking/ad free websites.

Keep track of list changes. So, if someone gets paid off to allow certain sites on their popular list, others can easily fork a past version of the list.

Re: We can do better than DuckDuckGo

#177
post #47
post #34

Earlier quoted context omitted.

Maybe instead of hard-coding these preferences in the search engine, or having it try to guess for you based on your search history, you can opt-in to download and apply such lists of ranking modifiers to your user profile. Those lists would be maintained by 3rd parties and users, just like eg. adblock blacklists and whitelists. For example, Python devs might maintain a list of search terms and associated urls that g…

I like this idea! I think the biggest difficulty with it - which is also probably the most important reason that engines like Google and DDG are currently struggling to return good results - is that the search space is just so enormously large now. The advantage of the suggestion in the blog post is that you trim down the possible results to a handful of "known good" sources. As I understand it, you'd want to continu…

This may be a very dumb question, but could the filtering be done client-side? As in, DDG's servers do their thing as normal and return the results, then code is executed on your machine to weight/prune the results according to your preferences.

Maybe this would require too much data to be sent to the client, compared to the usual case where it only needs a page of results at a time. If so, would a compromise be viable, whereby the client receives the top X results and filters those?

Re: We can do better than DuckDuckGo

#178

I’m pro privacy, but I dont have a problem with AdWords, outside of googles implementation. If AdWords targeting was purely based on the search term, I don’t mind. The search engine has to generate revenue somehow, and the revenue generated on “Saas crm” with a single click is likely to be larger than any users annual subscription. (10 - 100+ per click) I’m unclear on the ethical / privacy concerns of “AdWords” style…

Here is the problem (and it isn't privacy):

A search engine's job is to present you with the best possible results for any given query.

A ad is either A) the best possible result or B) not the best possible result. If the ad is the best possible result, then the search engine must display it anyway in order to fulfill its mission. If it is not the best possible result, the search engine must violate its mission in order to display it. To put it bluntly, advertising is paying to decrease the quality of search results.

Re: We can do better than DuckDuckGo

#179

> Instead, it should crawl a whitelist of domains, or “tier 1” domains. These would be the limited mainly to authoritative or high-quality sources for their respective specializations, and would be weighed upwards in search results. Not a big fan of this conclusion. Who chooses the white list, and why should I trust them? Is it democratically chosen? Just because a site is popular very clear does not mean it's trustw…

> If I want my blog to show up on your search engine, do I have to get it linked by one of those sites, or can I register with you? Will I be tier 1, or

I think what I'd say in defense is that we've misunderstood what search engines are useful for. They're really bad at helping us discover new things. Your blog might be awesome, but it's not going to be easy for a search engine to tell that it's awesome. It's going to have to compete with other blogs that also want views, some of whom are going to be better than yours at SEO, and so on.

What a search engine might be able to tell is that it's useful. Because what search engines are at least potentially good at is answering questions. You do that by having a list of known good sites to answer specific types of questions, and looking at the sites they link to. It's when you try to do both (index everything on the web and provide accurate answers to specific questions) that you end up failing to do either. For example this is the #2 result for "python f strings" on DDG[1]. It's total garbage, and, quoting the blog, "we can do better". (This result is also on page 1 for the same query on Google.)

What I believe ddevault is suggesting is that we make a search engine that does the only thing search engines are really good at, answering questions. You throw away the idea of indexing everything on the web, and therefore the possibility of "discovery". What that means is that in 2020 you need some other mechanism for discovering new sites, bloggers, and so on. Fortunately we do have some alternatives in that space.

To be clear, I don't know if I 100% buy this argument, but I think it's the general idea behind what's being suggested in this blog post.

[1] https://careerkarma.com/blog/python-f-string/

Re: We can do better than DuckDuckGo

#180
post #70
post #52

Earlier quoted context omitted.

Author didn't even DDG to find this out?

Drew has a longstanding history of ill-informed rants ([1] [2]) about technology. He's also quite willing to lie about the facts[3]. [1] https://news.ycombinator.com/item?id=24121609 [2] https://news.ycombinator.com/item?id=23966778 [3] https://news.ycombinator.com/item?id=24023998

No personal attacks on HN, please.

https://news.ycombinator.com/newsguidelines.html

Digging up past internet history as ammunition in an argument isn't cool in general. It's not that such details are necessarily wrong or irrelevant, but doing this has a systemically degrading effect and we don't want to be that sort of community.

https://hn.algolia.com/?dateRange=all&page=0&prefix=true&sor...

Post reply on HN