Live data from Hacker News

Show HN: Imagine a search engine that removed top million sites from its index

millionshort.com

81–90 of 204 posts

Re: Show HN: Imagine a search engine that removed top million sites from its index

#81
post #59

It means that our ranking algorithms have good recall but very poor precision. We value web page connectivity more than its content. We don't know how to teach machines to evaluate web page for its merit so we hope that a large number of Twitting, Liking and Plusing non-experts will approximate single expert. Millionshort shows that this model isn't good enough.

It shows the problem with using tags for ranking rather than considering the source of the information.

Yahoo became the goto portal because the quality of their links was high. Google overtook them because metadata was the only practical way to keep up with the pace at which the web was expanding.

However, after more than a decade of SEO, it's limitations are evident as the long tail gets increasingly harder to reach with search engines based on the Google model, e.g. "Purchase empiricist philosophers on eBay."

My suspicion is that Google is abandoning neutral search for personalized search in part because of the problems SEO represents when it comes to neutral searching and personal tracking provides an algorithmic way of establishing the quality of results for ranking purposes. One which is easier than attempting to develop a neutral curation algorithm.

Re: Show HN: Imagine a search engine that removed top million sites from its index

#82

Removing Wikipedia might be a mistake. Otherwise, it's great.

Links to Wikipedia could appear in a separate box near the top...it should be this way on Google and Bing as well...because Wikipedia is either exactly what someone wants or exactly what they don't want.

Re: Show HN: Imagine a search engine that removed top million sites from its index

#83
This is the first off-brand search engine I've seen that's, in some sense, cooler than Google.

For one thing, the huge miasma of spam websites that dominates the SERPs just isn't there -- I hope this lights a fire under Google's butt and people see another world is possible.

Re: Show HN: Imagine a search engine that removed top million sites from its index

#85

Removing Wikipedia might be a mistake. Otherwise, it's great.

Why would anyone want Wikipedia in search results is beyond me. If I want to read a Wikipedia article, I can just search Wikipedia itself. I know the kind of information it has, so there is no point in "ranking" it against other websites.

Re: Show HN: Imagine a search engine that removed top million sites from its index

#86

This is the first off-brand search engine I've seen that's, in some sense, cooler than Google. For one thing, the huge miasma of spam websites that dominates the SERPs just isn't there -- I hope this lights a fire under Google's butt and people see another world is possible.

Thanks for the great comparison.

Re: Show HN: Imagine a search engine that removed top million sites from its index

#87

This is the first off-brand search engine I've seen that's, in some sense, cooler than Google. For one thing, the huge miasma of spam websites that dominates the SERPs just isn't there -- I hope this lights a fire under Google's butt and people see another world is possible.

It reminds me of why I first moved to Google from Yahoo/Webcrawler/Altavista/etc in the first place.

Re: Show HN: Imagine a search engine that removed top million sites from its index

#88
post #46

Well that's one way to break out of the filter-bubble/echo-chamber I suppose. If only our best search technology was based on something better than a popularity contest :(

Popularity I think is important. But not at the expense of relevance. It's not a easy nut to crack.

Why? If I search, say, for a game review, I don't care whether it comes from a popular website or a blog no one reads. In fact, the topmost websites are more likely to be biased, since they try to appease everyone and they also have strong relationships with publishers. The blog no one reads is nearly guaranteed to be honest (if not well-written).

This holds true for most topics I can think of. Moreover, if I ever need to read Wikipedia and such, I already know about those websites, and I can go there directly - no need to search. Shouldn't web search engines act like discovery tools?

Re: Show HN: Imagine a search engine that removed top million sites from its index

#89
post #87

This is the first off-brand search engine I've seen that's, in some sense, cooler than Google. For one thing, the huge miasma of spam websites that dominates the SERPs just isn't there -- I hope this lights a fire under Google's butt and people see another world is possible.

It reminds me of why I first moved to Google from Yahoo/Webcrawler/Altavista/etc in the first place.

Yup. Since this is trivial for Google to copy, it's unlikely that it'll actually disrupt Google in any way, but still...very cool.

Re: Show HN: Imagine a search engine that removed top million sites from its index

#90
post #85

Removing Wikipedia might be a mistake. Otherwise, it's great.

Why would anyone want Wikipedia in search results is beyond me. If I want to read a Wikipedia article, I can just search Wikipedia itself. I know the kind of information it has, so there is no point in "ranking" it against other websites.

I'm not sure if this is still the case, but in the past Wikipedia's search engine was terrible and it was actually easier to google "X wikipedia" or "X wiki".
Post reply on HN