Live data from Hacker News

Google search and search engine spam

googleblog.blogspot.com

111–120 of 223 posts

Re: Google search and search engine spam

#111
post #83

Earlier quoted context omitted.

Assuming we are looking at the same results, the pages at position #6 and #10 are not copies of the stackoverflow content at position #8. They are copies of http://stackoverflow.com/questions/1399293/test-priorities-o... . Unfortunately, the only place that the word "delay" (which is in your query) sometimes appears on that stackoverflow page is in the "related" links in the right column. At the time Google last craw…

Yeah.. it might not be the best example, as what I was searching for is not current possible, so there are no correct results for it.

One UI issue we've struggled with is how to tell the user that there isn't a good result for their query. This comes up when we evaluate changes that remove crap pages all the time. For nearly any search you do, something will come up, just because our index is enormous. If the only thing in the result set that remotely matches the query intent is a nearly empty page on a scummy site, is that better or worse than having no remotely relevant results at all? I definitely lean towards it being worse, but many people disagree.

Re: Google search and search engine spam

#112

Curious though what the metrics they use to evaluate effectiveness against spam are. It could have just as much (or as little) spam indexed as it had 5 years ago, and in some comparisons that would be valid, but what if much of the spam had moved from being evenly distributed throughout results to being distributed in the top positions? Then, one could say spam was even lower than ever in total quantity, but it would…

I actually feel quite comfortable with our metrics. Back in 2003 or so, we had pretty primitive measures of webspam levels. But the case that you're wondering about (more spam, but in different positions) wouldn't slip past the current metrics.

How do you measure your own performance regarding spam?

If the measurement system is able to detect spam, why is spam not removed in the first place, before it has a chance to show up in the metrics...?

Re: Google search and search engine spam

#113
post #110

Earlier quoted context omitted.

I guess if you find it helpful, nothing. I always skip their results because they're always really shallow. The kind of article you'd expect if you were paying someone $4 to research and write an article on a topic they know nothing about.

They are content farms - plain and simple. You might consider throwing about.com in there, too. They're made to earn the employers revenue, not help out people with the particular queries.

Well, why not throw out anything ending in .com or .net? After all, they're all made to earn revenue...

I hope the point is clear -- "to earn revenue" isn't an appropriate test. But your point is well taken. I'd prefer a search to turn up the site made by the guy in the garage who is passionate about the subject rather than the $4/hr content farm content.

I think if we're to continue complaining about spam we have to actually define what we mean by spam. And I have this strange feeling that not everyone will agree...

Re: Google search and search engine spam

#114
post #98

While I applaud the direct personal response, I feel like the content says "we don't see a problem." If users see a problem and you don't, smaller competitors can eat your lunch. I'm kind of hoping for some competition in the field. In terms of adsense, if you really think about it, adsense content on a page should probably be a slightly negative ranking signal (not just not a positive signal). The very best quality…

I feel like the content says "we don't see a problem."

We see a virtually unbounded number of problems with our search results, and we're working constantly to fix them. Most of the people I talk to who work on search have the attitude that Google is horribly broken all the time, it's just also measurably the best thing available.

Google as a company, and search quality in particular, does not rest on its laurels. The people who hate Google's search results the most all work at Google. If you think you hate Google's search results as much as we do, you should come work for us. :)

Re: Google search and search engine spam

#115

Earlier quoted context omitted.

I've been tracking how often this happens over the last month. It's gotten much, much better, and one additional algorithmic change coming soon should help even more. I'm not saying that a clone will never be listed above SO, but it definitely happens less often compared to a several weeks ago.

My experience is the exact opposite: I am seeing many, many more clone sites in my search results in the last few months. It feels like it increases when I accidentally click a clone site. This happens for more than StackOverflow clones. Mailing lists, Linux man-pages, FAQs, published Linux articles, etc. all have clone pages that are obvious link farms (sometimes they even include ads that attempt to harm my compute…

Try DuckDuckGo. Gabriel has been doing an aggressive job about removing unsavory domains and I've been fairly impressed. I think that Google probably can't be nearly as aggressive for political reasons.

Re: Google search and search engine spam

#116

Earlier quoted context omitted.

Our policy in search-quality is not to pre-announce things, but we did give some pretty strong hints about planned improvements to search quality in that post (e.g. talking about scraper sites). I'll be happy to talk more about them soon when they launch.

Hey Matt, is there anything y'all can do about the content farm sites where someone buys an old high pr domain and sells 100's of links on it and drops them in between tons of content? Here's a prefect example of that: http://www.dcphpconference.com/ .

That is clearly webspam. We only indexed five pages from that site, but there's no reason for that site to show up at all. Thanks for mentioning it. It won't be in our index for much longer.

Re: Google search and search engine spam

#117

According to our metrics we are great; pity-about-you, unless you can "Please tell us how we can do a better job." I expected more. It reads like content farm.

Did you read past the first paragraph? They mention some recent changes they've made: To respond to that challenge, we recently launched a redesigned document-level classifier that makes it harder for spammy on-page content to rank highly. ... We’ve also radically improved our ability to detect hacked sites, which were a major source of spam in 2010. And we’re evaluating multiple changes that should help drive spam l…

[deleted]

Re: Google search and search engine spam

#118

According to our metrics we are great; pity-about-you, unless you can "Please tell us how we can do a better job." I expected more. It reads like content farm.

Did you read past the first paragraph? They mention some recent changes they've made: To respond to that challenge, we recently launched a redesigned document-level classifier that makes it harder for spammy on-page content to rank highly. ... We’ve also radically improved our ability to detect hacked sites, which were a major source of spam in 2010. And we’re evaluating multiple changes that should help drive spam l…

@ Matt, wondering if you can give any insight on how the new algo will affect syndicated news sites? Will it grade them as duplicate content and devalue them? And if so, will google news be held to the same crucible of truth?

Re: Google search and search engine spam

#119

"we’re evaluating multiple changes that should help drive spam levels even lower, including one change that primarily affects sites that copy others’ content and sites with low levels of original content." "As “pure webspam” has decreased over time, attention has shifted instead to “content farms,” which are sites with shallow or low-quality content. In 2010, we launched two major algorithmic changes focused on low-q…

I don't think they're talking about duplicate content so much as stuff like Demand Media (ehow), Associated Content, Hubspot, ezinearticles, etc, etc.

It'll be interesting to see whether Google will take on some of the bigger sites. Add BizRate to the list too (its pages have a 'Related Searches' keyword stuffing section, and 99% of the site is auto-generated), along with WiseGeek: possibly the worst and most spammy content farm I've seen, IMO.

I personally am always baffled how some of the really large, spammy sites with low quality content are rewarded with premium AdSense accounts and great rankings.

Re: Google search and search engine spam

#120
"The short answer is that according to the evaluation metrics that we’ve refined over more than a decade, Google’s search quality is better than it has ever been in terms of relevance, freshness and comprehensiveness."

The long answer is that without Wikipedia results, Google's search quality would be at an all-time low in terms of relevance, freshness and comprehensiveness.

Post reply on HN