Live data from Hacker News

We can do better than DuckDuckGo

drewdevault.com

331–340 of 383 posts

Re: We can do better than DuckDuckGo

#331
post #196

Earlier quoted context omitted.

This is the single suggestion that got me excited, instead of a behemoth do-it-all engine, have focused search engines. Exciting idea.

You may like 'boardreader.com'. I find it helps when looking for obscure info on random topics.

Wow, pretty cool! Thank you.

Re: We can do better than DuckDuckGo

#332

> The search results suck Do they really though, for normal people that is?. Some of my searches today below, can't remember the exact terms I used. Mix of DDG and Google. 1) Walt Whitman, I wanted a basic overview of his work to satisfy some idle curiosity. DDG gave me his wikipedia page. Bingo 2) EAN-13 check digit. First result wikipedia telling me how to calculate it. I see it is simple and I have a long list in…

It sucks for me when I search anything outside technical/science and daily life.

It also sucks at retrieving very new information.

And I say this as someone who set DDG as default.

I mean, you do seem to be DDG’s ideal user. You searched for mostly technical issues, and a hot political issue.

Re: We can do better than DuckDuckGo

#333

Earlier quoted context omitted.

I suspect that anyone who claims that Duckduckgo "Just works" only do english search. I usually do "english" / "mother tongue" searchs all day. Everytime, I need to remember to toggle the regional button otherwise I get attrocious results. Whereas google simply understand that if I'm searching using the english language it should prioritize english results while if I'm searching in another language it should prioriti…

For me (German) it’s different. With DDG, I can easily choose to search for German content (by using !ddgde), with google I have to hope that they search for what I want. Sometimes google does, sometimes it does not. And if it doesn’t I’m out of luck unless I go into the settings and look for a way to tell it what to do. Google automates, DDG leaves me to choose. I prefer the 2nd approach every time.

Thank you for the !ddgde bang - I face exactly the same problem as you.

Re: We can do better than DuckDuckGo

#334
One of the biggest problems the article points out more than anything is "Who's Going To Pay For It?"

You have one of two options. The crowd-funded approach would have to come with an understanding that you're trying to build a better, more private search engine that will leave the payer paying for all those who don't and the payer won't be able to have a say in anything as we'd like search results to be flat and even across the board. That means if I search for Coronavirus results, not only do I get the goverment and "Official" sources, I should get everything that I'm looking for and refine as needed.

The second approach is obviously big money but if you have big money coming in, big money will give You one of two options; Do as they say or they withdraw funding leaving you back at option one and having to downsize.

Rocks and hard places people. Not much else you can do there. Unless you take a Pilled.net approach.

Re: We can do better than DuckDuckGo

#335
post #215

Earlier quoted context omitted.

Instead of having many bots do inefficient crawling, web sites should publish their own index. Intermediate parties can combine indexes of the sites. Sites that do not provide indexes get less visitors.

This is the idea behind sitemaps which have existed forever.

Sitemaps are lists of urls on a site. They are not a text index.

Re: We can do better than DuckDuckGo

#337

Earlier quoted context omitted.

Aside from this a big reason to build this is it seems a lot simpler than writing a giant web crawler ala google and thus is a good target for an open source solution. Which is the biggest problem with duck duck go.

Do you have thoughts on implementing the distributed search? I'm thinking about playing around with this in my spare time, but that part seems the hardest to do.

Don't start from scratch, take a look at yacy (https://yacy.net/) that already does most of what is discussed here

Re: We can do better than DuckDuckGo

#338

Earlier quoted context omitted.

They literally user Bing api for search results which is well known

It's well-known that they use Bing. They also say that they use other sources. In this case I'm looking for an explicit disambiguation of what sources they use for what; my first read led me to interpret it as them also using their crawler to return search links (as opposed to just being used for their instant answers).

"We also of course have more traditional links in the search results, which we also source from multiple partners, though most commonly from Bing (and none from Google)."

This refers to the actual search results and if they used their bot for that I don't see why they would say "multiple partners" instead of "multiple partners and our own crawler". The fact that they don't have their own index is such a common "complaint", and this page is often referred, so if they really used their own bot they should have added that a long time ago.

And it's not just a legacy page they have forgot to update. It keeps being updated. In 2019 it said:

"We also of course have more traditional links in the search results, which we also source from a variety of partners, including Oath (formerly Yahoo) and Bing."

It's interesting to note that back in 2014 the page looked like this: https://web.archive.org/web/20131202065705/https://duck.co/h...

Here they talk about their own indexes getting bigger but at the same time admitting that "it seems silly to compete on crawling and, besides, we do not have the money to do so". Completely understandable but also interesting that the current page doesn't mention their own index at all. Maybe they used to have a goal to build their own independent index that has now been dropped?

All in all, I think it's safe to presume that their own crawler is only used for Instant Answers etc since that's the part of the sources where it's mentioned. Or at the very least used to such a small extent in the actual search results that it would be disingenuous to even mention it as a source.

Re: We can do better than DuckDuckGo

#340
post #215

> Instead, it should crawl a whitelist of domains, or “tier 1” domains. These would be the limited mainly to authoritative or high-quality sources for their respective specializations, and would be weighed upwards in search results. Not a big fan of this conclusion. Who chooses the white list, and why should I trust them? Is it democratically chosen? Just because a site is popular very clear does not mean it's trustw…

Instead of having many bots do inefficient crawling, web sites should publish their own index. Intermediate parties can combine indexes of the sites. Sites that do not provide indexes get less visitors.

That's the best way to fill your search engine with spam. There needs to be a third-party that verifies that the site-provided index is inline with the actual content. At which point said third-party can be a crawler.
Post reply on HN