Live data from Hacker News

We can do better than DuckDuckGo

drewdevault.com

211–220 of 383 posts

Re: We can do better than DuckDuckGo

#211
post #196
post #192

Earlier quoted context omitted.

This is not about a single search engine instance replacing all your search engine needs. There could be a community of Software developers running one instance of the OS search engine that focuses on programmers needs: documentation, vcs hosters, dev blogs, tech news websites and on topic blogs. Great if you need to search for how to solve a software issue, terrible if you need to figure out how long to cook spaghet…

This is the single suggestion that got me excited, instead of a behemoth do-it-all engine, have focused search engines. Exciting idea.

You may like 'boardreader.com'.

I find it helps when looking for obscure info on random topics.

Re: We can do better than DuckDuckGo

#212

> Instead, it should crawl a whitelist of domains, or “tier 1” domains. These would be the limited mainly to authoritative or high-quality sources for their respective specializations, and would be weighed upwards in search results. Not a big fan of this conclusion. Who chooses the white list, and why should I trust them? Is it democratically chosen? Just because a site is popular very clear does not mean it's trustw…

> If I want my blog to show up on your search engine, do I have to get it linked by one of those sites, or can I register with you? Will I be tier 1, or I think what I'd say in defense is that we've misunderstood what search engines are useful for. They're really bad at helping us discover new things. Your blog might be awesome, but it's not going to be easy for a search engine to tell that it's awesome. It's going t…

> You do that by having a list of known good sites to answer specific types of questions, and looking at the sites they link to.

I mean, that's basically the core of original Google pagerank, right? A "good" site linking to another site is what makes that other site some amount of "good" too, links from better sites carry more 'juice'. "good" is of course not just binary, but a quantitative weight.

I don't know to what extent that's still at the core of their relevancy rankings. I don't know how all those annoying spammy recipe blogs or content farms get to the top of the results either. I don't think it's because Google's engineers believe they are "good" results.

Relevancy ranking on web search is clearly a hard problem, mainly because so many authors are trying to game it, it's a feedback loop.

If Google can only do as google does despite pouring a whole lot of money into it, I don't see a reason to bet that better will be what's basically an over-simplified description of how Google started out doing it (and then evolved it because it wasn't good enough).

Re: We can do better than DuckDuckGo

#213
post #13

One thing I always wished for is if there were a way to use duckduckgo bang searches in my browser without sending them through DDG. But apparently it's harder to implement than it sounds.

Chromium implements this as a feature by default. Visiting a website with an OpenSearch tag in its `head` or searching on a website without one lets you later search the website itself from the urlbar by pressing "tab". It's history-based and works very well.

https://www.chromium.org/tab-to-search

Re: We can do better than DuckDuckGo

#214
post #131

Earlier quoted context omitted.

If you want a good engine there is no need to index 100B pages, since 99% of the pages are blogspam.

How are you going to identify what's blogspam and what's legitimate without indexing it all in the first place?

You load, detect and discard it?

We’re already talking about building a search engine, might as well talk about a model to convincingly detect blogspam too.

Re: We can do better than DuckDuckGo

#215

> Instead, it should crawl a whitelist of domains, or “tier 1” domains. These would be the limited mainly to authoritative or high-quality sources for their respective specializations, and would be weighed upwards in search results. Not a big fan of this conclusion. Who chooses the white list, and why should I trust them? Is it democratically chosen? Just because a site is popular very clear does not mean it's trustw…

Instead of having many bots do inefficient crawling, web sites should publish their own index. Intermediate parties can combine indexes of the sites. Sites that do not provide indexes get less visitors.

Re: We can do better than DuckDuckGo

#216
post #199
post #160

Earlier quoted context omitted.

Just don't give him IRC ops and then get into a private argument with him. https://www.omnimaga.org/news/omnomirc-moved-to-new-server/m...

Quite apart from this type of attack breaking the site guidelines and not being allowed on HN (about which see https://news.ycombinator.com/item?id=25130908 )... that was 9 years ago. Imagine being publicly shamed for the worst thing you've done in the past decade. I don't think anyone is going to pass that test. Is this the kind of world you (or any of us) really want to be part of? Surely not. Therefore please don'…

I think we also have the right to know as I wasn't aware of such past but now I can't see the page.

Re: We can do better than DuckDuckGo

#217

The way I'd code a better search engine is I'd design an ML model that's trained to recognize handwritten HTML like this, and only add those to the index. It'd be cheap to crawl probably only needing a single computer to run the whole search engine. It'd resurrect The Old Web, that still exists, but just got buried beneath the spammy SEO optimized grifter web over the years as normies flooded the scene.

I hope to never use your search engine. I love hand written HTML as much as the next guy, but search engine's are made to find things. And useful information exists on web sites that use generated and/or minified HTML.

Thanks for that buzzkill. I guess the lesson is if you can't do everything Google does, don't even try.

Re: We can do better than DuckDuckGo

#218
post #21

> they’ve demonstrated gross incompetence in privacy Not sure I buy the example that is given here. 1. It's an issue in their browser app, not their search service. 2. It's not completely indefensible: it allows fetching favicons (potentially) much faster, since they're cached, and they promise that the favicon service is 100% anonymous anyway. 3. They responded to user feedback and switched to fetching favicons loca…

Agreed. I think the key point here is that the web is a radically different place than it was in 1998 (when Google launched and established the search engine paradigm as we know it). Back then the quality-to-spam ratio was probably much higher, the overall size of the web was certainly much smaller (making scraping the entire thing more tractable), and there were many more self-hosted sources rather than platforms (m…

But the situation with authoritative domains hasn't changed much, and what "platforms" tend to be strong for answering questions? As in 1998, there are a few very good places for getting answers to certain kinds of questions. They are not facebook or twitter, ever.

Re: We can do better than DuckDuckGo

#220

> The search results suck Do they really though, for normal people that is?. Some of my searches today below, can't remember the exact terms I used. Mix of DDG and Google. 1) Walt Whitman, I wanted a basic overview of his work to satisfy some idle curiosity. DDG gave me his wikipedia page. Bingo 2) EAN-13 check digit. First result wikipedia telling me how to calculate it. I see it is simple and I have a long list in…

In consumer search there is a really long tail of questions (in 2017 15% of Google's daily queries have never been seen before[1]) and performance on this is very important.

I just searched for "lockdown rules for SA" (I'm in South Australia and we just had a new 20 person cluster, so we are going back into lockdown).

On DDG the first results was a Guardian article which was good, but then the rest were a mix of South African articles and blog spam. There were no SA Gov pages on the first page of results.

On Google the first result was the South Australian gov site with the rules, the second was the Guardian article, then more SA Gov pages and at result 8 I got a South African result.

https://searchengineland.com/google-reaffirms-15-searches-ne...

Post reply on HN