Live data from Hacker News

Google search and search engine spam

googleblog.blogspot.com

101–110 of 223 posts

Re: Google search and search engine spam

#101

Earlier quoted context omitted.

I've been tracking how often this happens over the last month. It's gotten much, much better, and one additional algorithmic change coming soon should help even more. I'm not saying that a clone will never be listed above SO, but it definitely happens less often compared to a several weeks ago.

Why not just downrank the results from the SO clones?

Because that wouldn't solve the problem for clones of other sites, or clones in other languages. And the Stack Overflow cloners could just make other websites. That's why a primary instinct in search quality is to look for an algorithmic solution that goes to the root of the problem. That approach works across different languages, sites, and if someone makes new sites.

To be clear: the webspam team does reserve the right to take manual action to correct spam problems, and we do. That not only helps Google be responsive, it also improves our algorithms because we get use that data to train better algorithms. With Stack Overflow, I especially wanted to see Google tackle this instance with algorithms first.

Re: Google search and search engine spam

#102
post #81

Earlier quoted context omitted.

It's safe to say that many Googlers read what people write on the web and talk about it internally, even if we don't always respond. We're power users too, so if a search result bothers you, it almost certainly bothers us too. :)

Matt, Can you speak about the possibility for personal domain blacklists for Google accounts? I know giving users the option to remove sites from their own search results is talked about a lot in these HN threads. Is there any talk internally about implementing something like this?

We've definitely discussed this. Our policy in search quality is not to pre-announce things before they launch. If we offer an experiment along those lines, I'll be among the first to show up here and let people know about it. :)

Re: Google search and search engine spam

#103
post #12

Earlier quoted context omitted.

Happened to me not ten minutes ago with the search string "pass json body to spring mvc" The efreedom answer at the 5th position is actually the most relevant - the stackoverflow question from which it was copied doesn't even show up on the first page. There is one stackoverflow result on the first page, but it deals with a more complex related issue, not the simple question I was looking for.

In Google's most recent cache, the efreedom result has the word "pass" on the page due to some related links content near the bottom, whereas the stackoverflow page does not. If you modify your query to [parse json body to spring mvc], stackoverflow is at position #1, and efreedom is at position #4. This still has room for improvement, but it would seem like the simplest explanation is just the better match on your q…

Didn't notice that - that's good to know. That's actually exactly how I'd expect a good search engine to behave. As annoyed as I am when I get a junk result, I'd be even more pissed if Google dropped terms from my query just so it can return a more popular site.

Of course, then all the content-copy farms will respond by copying valid content plus word lists - hopefully Google knows how to detect that.

Re: Google search and search engine spam

#104

Earlier quoted context omitted.

Why not just downrank the results from the SO clones?

Because that wouldn't solve the problem for clones of other sites, or clones in other languages. And the Stack Overflow cloners could just make other websites. That's why a primary instinct in search quality is to look for an algorithmic solution that goes to the root of the problem. That approach works across different languages, sites, and if someone makes new sites. To be clear: the webspam team does reserve the r…

But detecting duplicate content should not be very difficult, esp. now that Google indexes everything almost in real time. The site that had the content first is necessarily canonical and the others are the copies?

Because we don't understand what's hard, we think you're not really trying, and then we make up evil reasons to explain that.

I believe if people understood better the difficulties of spam fighting they would be more understanding.

Re: Google search and search engine spam

#105
post #4

Am I the only one who was really hoping for some specifics about what they're doing and plan to do about content farm rankings? Without that, the article is virtually devoid of content other than "we're really not so bad!" Edit: By specifics, I don't necessarily mean implementation details, just anything more informative and plan-of-action than acknowledging the problem.

Our policy in search-quality is not to pre-announce things, but we did give some pretty strong hints about planned improvements to search quality in that post (e.g. talking about scraper sites). I'll be happy to talk more about them soon when they launch.

Hey Matt, is there anything y'all can do about the content farm sites where someone buys an old high pr domain and sells 100's of links on it and drops them in between tons of content? Here's a prefect example of that: http://www.dcphpconference.com/.

Re: Google search and search engine spam

#106
post #5

Earlier quoted context omitted.

Can you provide a query where that is still the case ? It hasn't happened for a few weeks for me, since stack overflow changed their title seo.

What I find seriously bad is that even a huge site like stackoverflow has to optimize its search engine strategy to fight the problem. Little web sites are doomed.

SEO is supposed to be StackOverflow's core competency. They are completely aware most people end up on their site via Google. The search on their own site sucks.

The reason Q&A sites are so visible is that people tend to type questions in their search engines, so Q&A sites are a good match to those.

Re: Google search and search engine spam

#107

Earlier quoted context omitted.

Why not just downrank the results from the SO clones?

Because that wouldn't solve the problem for clones of other sites, or clones in other languages. And the Stack Overflow cloners could just make other websites. That's why a primary instinct in search quality is to look for an algorithmic solution that goes to the root of the problem. That approach works across different languages, sites, and if someone makes new sites. To be clear: the webspam team does reserve the r…

> Because that wouldn't solve the problem for clones of other sites, or clones in other languages.

Sites that are the victims of content cloning have to be very visible and valuable, so maybe a little manual curating could be relevant.

> the Stack Overflow cloners could just make other websites

Not really? The point is not to tag the clones but to tag the original; everything that is not the original and that has copied content is a clone -- its name, domain or country notwithstanding.

Re: Google search and search engine spam

#108
post #107

Earlier quoted context omitted.

Because that wouldn't solve the problem for clones of other sites, or clones in other languages. And the Stack Overflow cloners could just make other websites. That's why a primary instinct in search quality is to look for an algorithmic solution that goes to the root of the problem. That approach works across different languages, sites, and if someone makes new sites. To be clear: the webspam team does reserve the r…

> Because that wouldn't solve the problem for clones of other sites, or clones in other languages. Sites that are the victims of content cloning have to be very visible and valuable, so maybe a little manual curating could be relevant. > the Stack Overflow cloners could just make other websites Not really? The point is not to tag the clones but to tag the original; everything that is not the original and that has cop…

That would be a rather awesome feature for evil people: Just copy&paste your competitions content onto SO (or any other specially protected site) and their google ranking will drop like a stone.

Re: Google search and search engine spam

#109

In other industries, the government regulates that there must be a "Chinese Wall" (no communication or shared personel) between segments of a company that have conflicts of interest. In this case, search and advertising qualify. I expect to see proposals for this kind of regulation in the next 10 years as the FCC begins regulating the internet.

Good luck, see the Comcast & NBC merger. The FCC is impotent.

Re: Google search and search engine spam

#110
post #72

Earlier quoted context omitted.

What's wrong with eHow?

I guess if you find it helpful, nothing. I always skip their results because they're always really shallow. The kind of article you'd expect if you were paying someone $4 to research and write an article on a topic they know nothing about.

They are content farms - plain and simple. You might consider throwing about.com in there, too.

They're made to earn the employers revenue, not help out people with the particular queries.

Post reply on HN