SearchHut
11–20 of 161 posts
Re: SearchHut
#12>SearchHut indexes from a curated set of domains. Are there already plans to expand it as a service? E.g. subreddits could maintain their preferred lists of domains.
Re: SearchHut
#13I wonder if simply parsing all of Wikipedia (dumps are available, and so are parsers capable of handling them) and building a list of all domains used in external links would do the trick. Wikipedia already has substantial quality control mechanisms, and the resulting list should be essentially free of blog spam and other low-quality content. Wikipedia also maintains "official website" links for many topics, which can be obtained by parsing the infobox from the article.
Re: SearchHut
#14[citation needed]
The quality of the results right now are not very high, and in theory I don't understand why one would believe a search engine with a hand picked set of domains would be expected to outcompete a search engine that can crawl the entire web and determines reputation by itself. This also ignores the fact that a lot of domains have a mix of high quality content and low quality content, for example twitter or medium. If you are going to rely on domain-level reputation then your search engine is going to be way behind the search engines that can judge content more specifically, which is all of the other search engines.
If you were to tell me curated domains is just a bootstrapping method and as the search engine evolves it will change, fine, but right now the search engine is so simplistic that the theory of how it might be good is really the only point. And if that underlying theory is dubious, and the infrastructure is simplistic and obviously won't scale, then I don't know what is interesting or novel about this right now. Doesn't seem worthy of reaching the top of HN.
Re: SearchHut
#15Re: SearchHut
#16> SearchHut indexes from a curated set of domains. The quality of results is higher as a result, but the index covers a small subset of the web. [citation needed] The quality of the results right now are not very high, and in theory I don't understand why one would believe a search engine with a hand picked set of domains would be expected to outcompete a search engine that can crawl the entire web and determines rep…
Then why do Google and DuckDuckGo return 90% garbage for most queries?
"All of the other search engines" have completely failed to keep pages from the results that are not only low-quality, but outright spam.
Re: SearchHut
#17> SearchHut indexes from a curated set of domains. The quality of results is higher as a result, but the index covers a small subset of the web. [citation needed] The quality of the results right now are not very high, and in theory I don't understand why one would believe a search engine with a hand picked set of domains would be expected to outcompete a search engine that can crawl the entire web and determines rep…
Seems like your expectations are misplaced. Being at top of HN is not an indicator of quality, just interest.
Re: SearchHut
#18SearchHut: The first result is django which is not the most popular web server.
Google: Shows an answer box with the market share of various web servers.
Re: SearchHut
#19> SearchHut indexes from a curated set of domains. The quality of results is higher as a result, but the index covers a small subset of the web. [citation needed] The quality of the results right now are not very high, and in theory I don't understand why one would believe a search engine with a hand picked set of domains would be expected to outcompete a search engine that can crawl the entire web and determines rep…