Live data from Hacker News

Show HN: Imagine a search engine that removed top million sites from its index

millionshort.com

151–160 of 204 posts

Re: Show HN: Imagine a search engine that removed top million sites from its index

#151
You think that top results in Google and other commercial search engines are always ranked based on "popularity"?

It would be harsh to call this naive, but it shows a serious lack of SEM and SEO knowledge. Ever heard of "paid placement"?

Many years ago when Digital's AltaVista was our main search engine, it was becoming loaded down with paid placement.

The results were polluted.

Google eventually became the "clean" solution.

But now it's Google that is loaded down with all sorts of commercial crud, much of pointing to Google acquisitions.

And paid placement, among numerous other strategies, new and old, still exists.

The simplicity of millionshort is brilliant.

Filter out the crap.

Re: Show HN: Imagine a search engine that removed top million sites from its index

#152
Thankyou! I can see this being something I regularly use.

It may be a simple idea, but its something nobody else has done before, and I think the creators deserve a lot of credit for coming up with and implementing it. I hope they manage to get something from it. I can see that if the site becomes popular it will just get copied by other search sites.

Re: Show HN: Imagine a search engine that removed top million sites from its index

#153
"Quality" is subjective.

More relevant is _accuracy_, i.e., you get what you specify via search operators, and results are not influenced by all of Google's silly "factors". You know what you're looking for and how to frame the query. But Google assumes you're dumb and thinks it should decide for you.

Alexa Top 1M is a nice filter because the data comes from the Alexa Toolbar which only the most braindead web users would have installed. So you are in effect avoiding sites that the web's most braindead users would often visit.

Ranking sites based on "popularity" is great until you reach the point where the majority of users are not very intelligent. (cf. search engine users in 2004 with search engine users today.) When you reach that point, you get results where "quality" is determined by idiots (and SEO hats), not a group of intelligent peers.

Re: Show HN: Imagine a search engine that removed top million sites from its index

#155
Add a way for me to put this as my search engine in my firefox search bar.

Please.

EDIT: In trying to accomplish this task I found an add-on that lets you do this for anything.

(https://addons.mozilla.org/en-US/firefox/addon/add-to-search...)

Re: Show HN: Imagine a search engine that removed top million sites from its index

#157

Add a way for me to put this as my search engine in my firefox search bar. Please. EDIT: In trying to accomplish this task I found an add-on that lets you do this for anything. ( https://addons.mozilla.org/en-US/firefox/addon/add-to-search... )

Awesome.

Re: Show HN: Imagine a search engine that removed top million sites from its index

#158
post #116

It's a similar measure that's often used in NLP. Sentences, documents etc. are usually stripped of common or popular terms first and the remaining ones tend to have higher information value. It's not entirely a surprise that it works for meta-language constructs like the web and site popularity.

Uhh, 1) I am basically certain that it doesn't work. Imagine something this simple actually did work in general. Do you think Google, Bing, etc, wouldn't have implemented it? 2) I think the analogy is extremely flawed. nytimes.com is by no means a web page equivalent to "the" or "a" in language. Articles, pronouns, etc, don't really carry meaning. Despite being popular, nytimes certainly does.

I don't know. I haven't really spent a lot of time looking at this to get a good feel of the quality. But consider this, NYT tends to be a secondary source: e.g. book reviews (books are the primary source), science news (scientific papers are primary), business news (market movements and press releases are primary), etc.

Now consider, if I'm interested in some particular thing in science, is it better to get the NYT science reporting or just get the professor's publication and research page on their university website? Filtering out the top-n sites is more likely to turn up the professor's site near the top of the pile rather than after all of the popular sites' second and third level reporting.

Is this better? Depends on the audience. A "popular science" goal would argue the former is the better as science news simplifies, abstracts and popularizes complex science (with varying degrees of quality) while a scientist would prefer the latter.

Re: Show HN: Imagine a search engine that removed top million sites from its index

#159

I'm getting odd results with the following query : Search String : Ruby Remove From Top : 1000 & 10000 In both instances, the top hit is http://www.ruby-lang.org , which is also the top hit from both Google and DDG. Am I missing something? edit: formatting

I think that removes the top 1,000 or 10,000 across all searches, not just "ruby".

Re: Show HN: Imagine a search engine that removed top million sites from its index

#160

It's interesting to see popularity used as an inverse corollary with quality. Imagine a TV that skipped the most popular programming (goodbye American Idol), or a radio station that only plays non-hits. Of course, there are great websites out there that are very popular (Wikipedia, NYTimes/WSJ, StackOverflow). I'd love to see a search engine with a better signal for quality than non-popularity (this search engine), o…

Should be popular with hipsters.

I think you're probably right, but do you mean that as a "dig" against millionshort? You may be entering meta-meta-contrarianism: the hipster of hipsters. :)

http://lesswrong.com/lw/2pv/intellectual_hipsters_and_metaco...

Post reply on HN