Live data from Hacker News

We can do better than DuckDuckGo

drewdevault.com

261–270 of 383 posts

Re: We can do better than DuckDuckGo

#261
post #112
post #29

Earlier quoted context omitted.

"How do you save a search in a non-personally identifiable way?" To a first approximation, you just... do it. Granted, if you search "{jerf's realname here} {embarrassing disease} cure" or something, in the pathological case, you could at least guess that maybe it was me, though even then my real name is far from unique, and nothing stops anyone else from running such a search. But otherwise, if all you have is a pil…

At the very least your example is PII which you cannot save and also claim to be Private.

The mere existence of someone is not really PII. You don't know that I did that search, nor can you connect to anything else... and this is a constructed example in which I try to jam some sort of PII into a single search is itself a bizarre example that probably corresponds to fewer than 1 in 100,000 or 1 in 1,000,000 searches, if that. When's the last time you stuck your own PII into a search box and connected it to something of some sort of significance? It's a very small edge case.

A search history can reveal many things about a person. The mere fact that someone, somewhere searched for "star wars harry potter crossover slash", unconnected to any other search item, doesn't reveal anything about anybody.

Re: We can do better than DuckDuckGo

#262
post #63

DDG does operate their own crawler[1], though they also do still rely on third parties[2]. [1] https://help.duckduckgo.com/duckduckgo-help-pages/results/du... [2] https://help.duckduckgo.com/duckduckgo-help-pages/results/so...

Their own crawler is only used to fetch things for the widgets, not the search index.

Ahh, that does seem like a more correct reading of the page. Do we have a source that's unambiguous? Seems strange to have a bot that's able to parse pages for their instant answers but then not use those same results for regular search.

Re: We can do better than DuckDuckGo

#263

> Instead, it should crawl a whitelist of domains, or “tier 1” domains. These would be the limited mainly to authoritative or high-quality sources for their respective specializations, and would be weighed upwards in search results. Not a big fan of this conclusion. Who chooses the white list, and why should I trust them? Is it democratically chosen? Just because a site is popular very clear does not mean it's trustw…

Things actually sort of ran that way once. The DMOZ directory system was a cannonical, list of sites by subject listed top-down. It was maintained by a community of volunteers in the fashion of Wikipeda. I believe it was used as one reference for Google and other search engines at one time. I don't know if such an "objective" system could be rebuilt, however.

Still, it's good to remember that it was once uncertain whether people should access the web using something a table of content (portal/directory) or something like an index (search engine). It seems the search engines won.

See: https://en.wikipedia.org/wiki/DMOZ

Re: We can do better than DuckDuckGo

#265
> If SourceHut eventually grows in revenue — at least 5-10× its present revenue — I intend to sponsor this as a public benefit project, with no plans for generating revenue. I am not aware of any monetization approach for a search engine which squares with my ethics and doesn’t fundamentally undermine the mission. So, if no one else has figured it out by the time we have the resources to take it on, we’ll do it.

Now _that_ is putting your money where your mouth is!

Glad to see a technology leader taking this important issue head-on.

Re: We can do better than DuckDuckGo

#266

Aren't the secrets of the algorithm what prevent people from gaming the results? While I love the idea of search becoming fully open source I'm skeptical it could be done. I hope I'm wrong and I'd love to dedicate time to an open source project with this goal if anyone presents a convincing plan.

The algorithms, code and configuration can be public, if the ranking isn't done just by those, but instead by many participants in the project all over the world and also by personal preferences of each client. That would be hard to game.

Re: We can do better than DuckDuckGo

#267

Drew in his blog post talking about DuckDuckGo privacy issues, but his commercial startup Sourcehut does not offer the Privacy basics: 1. Account deletion 2. GDPR data request 3. Option to unsubscribe from emails So right now his blog reminds me one famous US politician Twitter account. Never fix your own problems, just blame others more often.

I normally would not reply to someone who equates my blog posts with the ravings of a megalomaniacal fachist, but I will at least clarify for the benefit of onlookers that all three of these points are false. I handle account deletion and GDPR requests all the time, and every email you get from sr.ht (1) is not a marketing email and (2) can be trivially unsubsribed from, with the exception of payment notifications -…

[deleted]

Re: We can do better than DuckDuckGo

#268
post #220

Earlier quoted context omitted.

In consumer search there is a really long tail of questions (in 2017 15% of Google's daily queries have never been seen before[1]) and performance on this is very important. I just searched for "lockdown rules for SA" (I'm in South Australia and we just had a new 20 person cluster, so we are going back into lockdown). On DDG the first results was a Guardian article which was good, but then the rest were a mix of Sout…

I got this: https://www.enca.com/news/sa-lockdown-strick-international-t...

Which is about South Africa - which might be a good result for you, depending on where you live.

So at least they are trying to to location based results.

Re: We can do better than DuckDuckGo

#269

> The search results suck Do they really though, for normal people that is?. Some of my searches today below, can't remember the exact terms I used. Mix of DDG and Google. 1) Walt Whitman, I wanted a basic overview of his work to satisfy some idle curiosity. DDG gave me his wikipedia page. Bingo 2) EAN-13 check digit. First result wikipedia telling me how to calculate it. I see it is simple and I have a long list in…

I’ve tried searching for stuff related to Alexa Presentation Language (APL) a bunch of times. It never finds anything useful; I throw “!g” on the query string and what I’m looking for is typically the first or second result.

Re: We can do better than DuckDuckGo

#270
post #220

> The search results suck Do they really though, for normal people that is?. Some of my searches today below, can't remember the exact terms I used. Mix of DDG and Google. 1) Walt Whitman, I wanted a basic overview of his work to satisfy some idle curiosity. DDG gave me his wikipedia page. Bingo 2) EAN-13 check digit. First result wikipedia telling me how to calculate it. I see it is simple and I have a long list in…

In consumer search there is a really long tail of questions (in 2017 15% of Google's daily queries have never been seen before[1]) and performance on this is very important. I just searched for "lockdown rules for SA" (I'm in South Australia and we just had a new 20 person cluster, so we are going back into lockdown). On DDG the first results was a Guardian article which was good, but then the rest were a mix of Sout…

Hrm, I think it's extremely iffy to abbreviate South Australia like that in a search query. You don't need the "for" either.

BTW, when I perform the same search, Google's first result is "What Are the Lockdown Rules for South Africa? A Guide for ..." and all the other results on the first page are about South Africa too. (Note: I'm in Japan)

Post reply on HN