Live data from Hacker News

We can do better than DuckDuckGo

drewdevault.com

141–150 of 383 posts

Re: We can do better than DuckDuckGo

#141

Earlier quoted context omitted.

I agree and I have been hoping Apple builds a serious competitor. I welcome any competition at this point. Let's be real, not many people are using bing. People _would_ actually use apple search.

>I welcome any competition at this point. Microsoft tried and failed to build a competitor and it's not like they have shallow pockets. They grossly underestimated a number of aspects: - The huge number of man-years invested in Google's search quality stack hand-tuning and what it would take to replicate it. - The fact that the machine learning field was simply not ready to tackle the search quality problem. - The in…

Sure but Microsoft has tried and failed to build a competitor to google not to DDG.

I don't see it that hopeless. I feel it kinda is like starting Open Street Maps. It won't be perfect for a long time but there will be people who'd prefer it and help out.

Re: We can do better than DuckDuckGo

#142
While I agree with its lack of organization I don't think YaCy being untolerably slow is necessarily an argument. If you are looking for a complete set of pages on a specific topic time is sort-of irrelevant. Google for example has alerts for new results. That these pages are not available sooner (before publication) is not intolerable. You can also throw hardware at YaCy and adjust the settings which improves it a lot. The challenge with a distributed approach is sorting the results. Other crawlers have the same problem but in a distributed system it is even harder.

Running an instance for websites related to your occupation or hobby YaCy is quite wonderful. You don't want google removing a bunch of pages that might cover exactly the sub-topic you are looking for. Of course the smaller the number of pages in your niche the better it works.

Re: We can do better than DuckDuckGo

#143
post #122

It's almost impossible to build a decent web search engine from scratch today (i.e., build your own index, fight SEO spams, tweak search result relevance...). The web is already so big and so complex. Otherwise Google won't need to hire so many people to work on search alone. If you didn't start at the very early stage of tiny web (e.g., Google in 1996 as a research project) and grew with the web over the past 20+ ye…

You are probably right. But still... the approach suggested makes kind of sense. A curated list of trusted sites as kind of seed. Not the entire web. This can be as small or as large as can be useful. It does not need to be about the entire web. How big is the "useful" blogosphere, for example? Cannot an opensource project that gathers momentum somehow create a curated list of let's say 10 000 trusted blogs and index those? Index all mailing lists that can be found, index all of reddit, index Hacker News, index Wikipedia, the 100 most well regarded news sites in each country, etc. Would not such an index be a good start and better than Google in many cases?

Re: We can do better than DuckDuckGo

#144
post #60

SEO is crushing the utility of Google. It is pretty telling when you need to add things like site:reddit.com to get anything of value. Harnessing real user experiences (blogs, etc) is the key to a better search engine. This model unfortunately crumbles under walled gardens which is increasingly the preferred location of user activity.

That’s where blogs were at, but now a massive portion of them are content farms / splogs.

You’re right that the walled gardens have hurt this. So often I search something specific, or a topic, and find very little. But I know there are communities on Facebook for this, I know there would be peoples posts out there on Instagram which 100% answer my question. But they may as well not exist. Unless I was “following” then when it was said, and mentally indexed it, these things are mostly unfindable, and that’s if I even have an account for said service (which I don’t for Facebook)

It’s sad, more people than ever using the internet, more content & knowledge being created than ever before, yet it’s no longer possible to find the great answers.

Re: We can do better than DuckDuckGo

#145

Why couldn't several coordinating specialized search engines share their data via something like "charge the downloader" S3 buckets? Then you get an org like StackExchange who could provide indexed data from their site and the algorithms to search the data the most efficiently, GitHub can do the same for their specific zone of speciality, Amazon, etc. Then anyone who wants to use the data can either copy it to their…

There's Common Crawl for the crawling aspect, about 3.2 billion pages last time I looked. One of the issues with that kind of detachment of jobs is crawl data freshness.

Re: We can do better than DuckDuckGo

#147
post #34

Earlier quoted context omitted.

Maybe instead of hard-coding these preferences in the search engine, or having it try to guess for you based on your search history, you can opt-in to download and apply such lists of ranking modifiers to your user profile. Those lists would be maintained by 3rd parties and users, just like eg. adblock blacklists and whitelists. For example, Python devs might maintain a list of search terms and associated urls that g…

This doesn't really seem immune from spam. I got signed up for goodreads (book review site), and I get tons of spam. It's not quite the same as your idea, but it is a currated list. I don't know how you stop spammers from adding bogus links in the python interest list (to use an example). This is a hard problem.. EDIT: Clarified goodreads reference!

Like any other list, it depends on who maintains it. You basically want to find the correct BDFL to maintain a list, much like many awesome-* repositories operate.

Re: We can do better than DuckDuckGo

#148

How would anybody ever know what the server is running and/or doing with the data you send it, regardless of if it is running open or closed source code? A service, running on somebody else's machine, is essntially closed. I think the only way to have an 'open' service is to have it managed like a co-op, where the users all have access to deployment logs or other such transparency. Even then, it requires implicit tru…

> How would anybody ever know what the server is running and/or doing with the data you send it, regardless of if it is running open or closed source code?

https://en.wikipedia.org/wiki/Homomorphic_encryption

Re: We can do better than DuckDuckGo

#149
My biggest pet peeve with DDG at the moment is that whenever I search for something on my phone the first two results are ads, and those two results actually take up my whole screen. I mean sure, those are probably not privacy invading, but I literally don't care as I wasn't looking for them.

Re: We can do better than DuckDuckGo

#150

Am I the only person who just doesn't have problems with DDG search results? What am I doing wrong (or right), here? I put a thing in and find it. I just don't use Google any more. Genuinely curious why it's working for me and such garbage for everyone else.

I sometimes come across inappropriate results - for example I search for a hex error code and the results are for other numbers - and sometimes the adverts are misleading, but neither are so prevalent enough that it harms the experience in general.

I always send feedback when I come across incorrect results and also try to when I get a really easy find.

I have not had to resort to any other search engine for at least five years.

Post reply on HN