Live data from Hacker News

Google search and search engine spam

googleblog.blogspot.com

121–130 of 223 posts

Re: Google search and search engine spam

#121
post #55

Earlier quoted context omitted.

I actually feel quite comfortable with our metrics. Back in 2003 or so, we had pretty primitive measures of webspam levels. But the case that you're wondering about (more spam, but in different positions) wouldn't slip past the current metrics.

How do you interpret the backlash from the users recently ? In your eyes, have we become more used to "perfect" results, or are the fewer bad results left more insidious and thus more harmful (despite the overall level of quality being higher) ? Personally I tend to find what I'm looking for by adding a few more words, but in the case of reviews and tech stuff it doesn't always work and I often have to rewrite my que…

in the case of reviews and tech stuff it doesn't always work

This is one of my pet peeves. If I search for "product X review", most of the result I get back are of the form "be the first to review product X", which is absolutely not what I want.

Re: Google search and search engine spam

#122

According to our metrics we are great; pity-about-you, unless you can "Please tell us how we can do a better job." I expected more. It reads like content farm.

Did you read past the first paragraph? They mention some recent changes they've made: To respond to that challenge, we recently launched a redesigned document-level classifier that makes it harder for spammy on-page content to rank highly. ... We’ve also radically improved our ability to detect hacked sites, which were a major source of spam in 2010. And we’re evaluating multiple changes that should help drive spam l…

To respond to that challenge, we recently launched a redesigned document-level classifier that makes it harder for spammy on-page content to rank highly

When is recently?

We’ve also radically improved our ability to detect hacked sites

Since when; all the complaints I have read are from the last few weeks, and when I looked up a medical complaint this week?

And we’re evaluating multiple changes that should help drive spam levels even lower

Oh, are you now? Of course... Translation 'We are looking into it'. D'uh.

that copy others’ content and sites with low levels of original content

This last part is the only datum I got from the article - they explicitly respond to stack overflow, etc... It is still fluff though.

What I would have preferred: I made a change on 2011-01-15 and you should see it here, here, and here. 'Here' can be broadly defined.

Re: Google search and search engine spam

#123
post #47

One misconception that we’ve seen in the last few weeks is the idea that Google doesn’t take as strong action on spammy content in our index if those sites are serving Google ads. That's not quite what I've been reading. I believe the more common claim is that Google has a disincentive to algorithmically weed out the kind of drivel that exists for no other reason than to make its publisher money via AdSense. It's abo…

I agree. Also interesting to see that Google defines webspam as "pages that cheat" or "violate search engine quality guidelines." By this definition, scraper sites are not spam at all. Nor are the spammy sites in my field which super-optimize for keywords in ways that make it difficult for legitimate content to rise to visibility. If Google did not operate AdSense, it seems hard to believe the company would not have…

Smaller competitors can't eat their lunch in web search, because all the content that was on the web is now on Wikipedia, YouTube, or Google Maps. Personally, I search these directly from the address bar. For the past four years I've only had two use cases for web search: 1. as a spell checker for proper nouns (and before Alpha, as a calculator) 2. to circumvent paywalls on scholarly papers by doing filetype:pdf on the title (works better than Scholar most of the time).

Re: Google search and search engine spam

#124
post #79
post #71

Earlier quoted context omitted.

While I'm not a fan of that site either, it's not true -- scroll to the bottom of the page. Sneaky? Absolutely, but the content and solution is there.

It's no longer true. Short version is: they used to, and got busted for, serving answers to the spiders and ads and pitches to the surfers. So now they show the answer at the bottom of a pile of ads and pitches. But they still suck. Horribly. And are the number one example I hear when people say "I wish Google would let me blacklist domains".

I don't believe no one has created a FF plugin for expert sex change (yet)!!

Edit: Even a GM script to remove all the leading spammy divs would do...

Re: Google search and search engine spam

#125
post #69

Google AdSense for Domains ( http://www.google.com/domainpark/ ) really makes a lie of not wanting useless content. The designed a revenue source for parkers/squatters.

+1 and my perfectly non-spammy blog was rejected by them. They didn't mention the reason but I have a hunch that it was because it isn't about any one or two things. It's about a bunch of things I find interesting - both personally as well as professionally. Now, google must have found it difficult to determine what ads to show, so they just black-listed my blog for ad-puposes. Period.

Re: Google search and search engine spam

#126

It doesn't appear that Cutts & Co. are looking to address any of the more popular blackhat link building methods that all popular SEO bloggers continually say "work but you shouldn't use them yourself because they're bad". Until the keyword "buy viagra" isn't littered with forum link and comment spam and parasite pages, Google's algo is still not "fixed"

People can make tens of thousands of dollars a month if they rank highly for certain phrases, so tons of SEOs (and spammers) are trying to rank for phrases like that. Other search engines might hard-code the results for [buy viagra], but Google's first instinct is to use algorithms in these cases, and if you pay attention, the results for queries like that can fluctuate a lot. With a billion searches a day, I won't c…

How about the more general case of gray and black hat link building techniques? Especially in internet marketing and certain SEO communities, everything from forum profile backlinking to blog comments and mass article submissions are done daily to rank sites in all niches, not just the really spammy viagra-type sites.

There's a growing concern/consensus that in loads of non-spammy niches, the only way to get decent rankings is to build links.

I know this is naturally something that would be very hard for Google to tackle (since if 'junk' backlinks = domain penalty, black hatters would simply spam their competitors' sites with backlinks and get their competitors deindexed), but are there active efforts being made so that gray and black hat link building campaigns aren't (in some cases) pretty essential to a site getting good rankings?

Re: Google search and search engine spam

#127
This is just lip service - Google's quality has dropped off and its really obvious. Recently, I've been regularly comparing the results I get from Google and the ones I get from Bing. Needless to say, Bing's far more relevant. The biggest point people have made are on the money searches like "MP3 Player" but the results I've been comparing have been local searches and programming things like: "show/hide text boxes in Javascript." In Google all I get is links to Amazon and other random results to link farms like javascriptworld.com. In Bing, I get links to forums and tutorials which is what I'm looking for.

Time and time again Google has failed. I've already moved on to Bing and Duckduckgo and I would recommend you do too. Unless you like digging through hordes of useless SERPS.

Re: Google search and search engine spam

#128
post #25

What are people's thoughts on companies who do create content farms? From the perspective as being a successful company rather than "I hate the spam and I hope they all DIAF". Personally I think any type of "scheming" in technology will eventually get caught and then all of a sudden there goes your business model.

It depends on the quality of the content, in my opinion. Ultimately magazines are 'content farms'. Various game review sites are 'content farms', since they're both designed to 'churn out' content and articles. (To give a couple of quick - albeit silly - examples). The difference is that the content in these two cases is high quality, usually containing images and possibly videos, and with lots of unique ideas and opinions given.

So I guess it depends. If a 'content farm' produces good, helpful content, then that's great and should be encouraged (even if it is done on a massive scale). But when it comes to a case like WiseGeek where content is spat out en masse, even if the content is crappy and really short, then it becomes a problem.

Re: Google search and search engine spam

#129
post #104

Earlier quoted context omitted.

Because that wouldn't solve the problem for clones of other sites, or clones in other languages. And the Stack Overflow cloners could just make other websites. That's why a primary instinct in search quality is to look for an algorithmic solution that goes to the root of the problem. That approach works across different languages, sites, and if someone makes new sites. To be clear: the webspam team does reserve the r…

But detecting duplicate content should not be very difficult, esp. now that Google indexes everything almost in real time. The site that had the content first is necessarily canonical and the others are the copies? Because we don't understand what's hard, we think you're not really trying, and then we make up evil reasons to explain that. I believe if people understood better the difficulties of spam fighting they wo…

> But detecting duplicate content should not be very difficult, esp. now that Google indexes everything almost in real time. The site that had the content first is necessarily canonical and the others are the copies?

Not necessarily. The rate at which Google refreshes its crawl of a site, and how deep it crawl, depend on how often a site updates and its PageRank numbers. If a scraper site updates more often and has higher PR than the sites it's scraping, Google will be more likely to find the content there than at its source. Identifying the scraper copy as canonical because it was encountered first would be wrong.

Re: Google search and search engine spam

#130
post #3

Metrics-shmetrics. Once I stop seeing StackOverflow clones listed above StackOverflow's original pages I will gladly believe that Google's search quality is "better than ever before."

I've been tracking how often this happens over the last month. It's gotten much, much better, and one additional algorithmic change coming soon should help even more. I'm not saying that a clone will never be listed above SO, but it definitely happens less often compared to a several weeks ago.

Have you noticed sites that scrape google groups content ranking higher than the google group they've scraped? How can that possibly happen? Still seems rampant.
Post reply on HN