Live data from Hacker News

We can do better than DuckDuckGo

drewdevault.com

221–230 of 383 posts

Re: We can do better than DuckDuckGo

#221
post #199
post #160

Earlier quoted context omitted.

Just don't give him IRC ops and then get into a private argument with him. https://www.omnimaga.org/news/omnomirc-moved-to-new-server/m...

Quite apart from this type of attack breaking the site guidelines and not being allowed on HN (about which see https://news.ycombinator.com/item?id=25130908 )... that was 9 years ago. Imagine being publicly shamed for the worst thing you've done in the past decade. I don't think anyone is going to pass that test. Is this the kind of world you (or any of us) really want to be part of? Surely not. Therefore please don'…

If I had done something like that so publicly in the past, and hadn't reached out to the people I hurt to reconcile. I would want people to publicly shame me for it. That way I can attempt to resolve it so everyone can have proper closure.

That said, message heard. I will refrain from this in the future here.

Re: We can do better than DuckDuckGo

#222

> The search results suck Do they really though, for normal people that is?. Some of my searches today below, can't remember the exact terms I used. Mix of DDG and Google. 1) Walt Whitman, I wanted a basic overview of his work to satisfy some idle curiosity. DDG gave me his wikipedia page. Bingo 2) EAN-13 check digit. First result wikipedia telling me how to calculate it. I see it is simple and I have a long list in…

I suspect that anyone who claims that Duckduckgo "Just works" only do english search. I usually do "english" / "mother tongue" searchs all day. Everytime, I need to remember to toggle the regional button otherwise I get attrocious results. Whereas google simply understand that if I'm searching using the english language it should prioritize english results while if I'm searching in another language it should prioritize it instead.

It gets tiring quickly and I find easier to append !g instead of clicking the regional toggle button.

Re: We can do better than DuckDuckGo

#223
post #215

> Instead, it should crawl a whitelist of domains, or “tier 1” domains. These would be the limited mainly to authoritative or high-quality sources for their respective specializations, and would be weighed upwards in search results. Not a big fan of this conclusion. Who chooses the white list, and why should I trust them? Is it democratically chosen? Just because a site is popular very clear does not mean it's trustw…

Instead of having many bots do inefficient crawling, web sites should publish their own index. Intermediate parties can combine indexes of the sites. Sites that do not provide indexes get less visitors.

This is the idea behind sitemaps which have existed forever.

Re: We can do better than DuckDuckGo

#224

Earlier quoted context omitted.

> to guess for you based on your search history, you can opt-in to download and apply such lists of ranking modifiers to your user profile pro-privacy does not sit well with terms such as search history and user profile

You might have misread. My proposal is an alternative to inferring user preferences based on their search history.

any type of profiling, opt-in or not, may be used to identify users

Re: We can do better than DuckDuckGo

#225

> The search results suck Do they really though, for normal people that is?. Some of my searches today below, can't remember the exact terms I used. Mix of DDG and Google. 1) Walt Whitman, I wanted a basic overview of his work to satisfy some idle curiosity. DDG gave me his wikipedia page. Bingo 2) EAN-13 check digit. First result wikipedia telling me how to calculate it. I see it is simple and I have a long list in…

I use ddg often myself.

Google does infer purpose better, and if someone is looking to buy something, it does well there too.

Ddg is very good at info queries and the more one uses it, the better it is.

What they could do is exactly what google did and that's to review those uses and improve.

But what they have right now is solid, given just a tiny bit of work.

Re: We can do better than DuckDuckGo

#226

> The search results suck Do they really though, for normal people that is?. Some of my searches today below, can't remember the exact terms I used. Mix of DDG and Google. 1) Walt Whitman, I wanted a basic overview of his work to satisfy some idle curiosity. DDG gave me his wikipedia page. Bingo 2) EAN-13 check digit. First result wikipedia telling me how to calculate it. I see it is simple and I have a long list in…

Yeah, all of these are quite DDG-friendly searches. It is my default engine and, yes, some results do suck quite consistently.

I'm a bit lazy right now to remember all the problems it has, but some of the most obvious are looking up for news on recent events (especially something small, stuff that doesn't appear in reuters and these sorts of media) and trying to find out some basic stuff about local shops and such (of course, I only know about how it feels in my location, not worldwide). On both occasions I pretty much always use "!g ..." right away, because DDG is just clueless about this shit. Google does this just fine (in fact, sometimes it's even impressive: there are thousands of cities like mine, yet Google can often tell me where I can buy some stuff I'd have no idea where to look for).

Re: We can do better than DuckDuckGo

#227

> Instead, it should crawl a whitelist of domains, or “tier 1” domains. These would be the limited mainly to authoritative or high-quality sources for their respective specializations, and would be weighed upwards in search results. Not a big fan of this conclusion. Who chooses the white list, and why should I trust them? Is it democratically chosen? Just because a site is popular very clear does not mean it's trustw…

> If I want my blog to show up on your search engine, do I have to get it linked by one of those sites, or can I register with you? Will I be tier 1, or I think what I'd say in defense is that we've misunderstood what search engines are useful for. They're really bad at helping us discover new things. Your blog might be awesome, but it's not going to be easy for a search engine to tell that it's awesome. It's going t…

As an experiment, I searched “tech news aggregator” on both google and DDG. Neither listed Hacker News. Instead, apart from a few actual sites, most of the links were articles saying “top ten tech news sites” or links to quora q&a threads.

It definitely seems that search engines can’t find new websites for people. Now they are just aggregating Q&A.

Re: We can do better than DuckDuckGo

#228
Low quality information is more profitable to produce than high quality information (thanks to Google and their ads). So all the incentives for people currently are to produce content that is nearly indistinguishable from spam. It is much more profitable to take 1 hour to write generic content than to take weeks to really think through everything. This problem will continue to get worse with AI generated text, and I don't see how Google can fight that. This is why 90% of my time is now spent on Youtube instead, which has the advantage that focusing on quality has a much higher ROI with video than with text. That doesn't mean it's not filled with garbage though, just that it's easier to swift through it. Product search is also something I now do on Amazon because Google results are essentially spam. The results there are also not ideal, but at least I can distinguish the bad results faster.

One way to make search better would be to embrace bias and give people what they want. Just accept that most information is biased and bring it to the forefront. Initially there would be some default domain whitelisting, but users can request sites to add or remove from their bubble. Maybe they can "share" these lists among each other and certain clusters would form, which you can then use to recommend more nodes. Users would also be clearly told that results are biased based on their preferences, and that they can choose to view results from other points of reference. It would also have different "modes" for things like products, information, and news. Maybe some users only want to see results from independent or ad-free sources, and they can choose those clusters. Maybe they want only far-right or far-left sources, and then they can choose those. At least people would be aware of their bias, which I think will help fight it more than acting like it doesn't exist. Essentially there would be an explorable graph so I can see different realities. It would be a combination of search and social media.

Re: We can do better than DuckDuckGo

#229
post #220

> The search results suck Do they really though, for normal people that is?. Some of my searches today below, can't remember the exact terms I used. Mix of DDG and Google. 1) Walt Whitman, I wanted a basic overview of his work to satisfy some idle curiosity. DDG gave me his wikipedia page. Bingo 2) EAN-13 check digit. First result wikipedia telling me how to calculate it. I see it is simple and I have a long list in…

In consumer search there is a really long tail of questions (in 2017 15% of Google's daily queries have never been seen before[1]) and performance on this is very important. I just searched for "lockdown rules for SA" (I'm in South Australia and we just had a new 20 person cluster, so we are going back into lockdown). On DDG the first results was a Guardian article which was good, but then the rest were a mix of Sout…

I got this:

https://www.enca.com/news/sa-lockdown-strick-international-t...

Re: We can do better than DuckDuckGo

#230

Why couldn't several coordinating specialized search engines share their data via something like "charge the downloader" S3 buckets? Then you get an org like StackExchange who could provide indexed data from their site and the algorithms to search the data the most efficiently, GitHub can do the same for their specific zone of speciality, Amazon, etc. Then anyone who wants to use the data can either copy it to their…

So you're proposing Snowflake for search?

It appears to be the case in a technical workflow sense, from the little I just read of Snowflake, but my proposal would be a much more open system than one under control of a single vendor. It'd be more like a set of standards for interoperation and a common data center so that the data is accessible under one roof. Maybe each entity could do specialized search as a service and the search aggregators would pay by providing some infrastructure to run the crawlers.

I don't personally think any system specification is impossible unless it goes against since mathematical law, so a really fast, distributed query system where there are a few hundred specialized providers for a single query is feasible. Imagine the aggregator does initial analysis to determine the category of search, like programming, news, or restaurant reviews, then sends the users query to a set of specialized providers that supply an index for that category, then fuse the results with some further analysis of the metadata returned. Then the user can also include or exclude the specialized providers at will.

You could also eliminate the aggregator as a service as simply make it a user application on the desktop, allowing for even more user control and maybe caching or something.

Post reply on HN