Live data from Hacker News

Matt Cutts is looking for scraper sites

twitter.com

81–90 of 120 posts

Re: Matt Cutts is looking for scraper sites

#81
post #46

Earlier quoted context omitted.

You won't find it on Google :) We're making a products recommendation site, focusing on goal-oriented decisions. We'll make the full announcement within the next couple of weeks. This reminds me. Google is really killing the "release early and release often" approach, if people will now have to do a ton of SEO learning and tweaking to avoid having your MVP permanently banned at launch day.

"Google is really killing the "release early and release often" approach, if people will now have to do a ton of SEO learning and tweaking to avoid having your MVP permanently banned at launch day." Judging by the downvotes this is not a legitimate concern? Why?

Actually, from an SEO perspective, the #1 principle Google follows is preventing "release often" from being effective for SEO.

Talk to anyone who makes money at PPC and they will tell you one thing. You make a campaign, measure the results, change it a little, measure the results, and make incremental improvements to make a profitable campaign.

If you could do that with SEO, SEO would be a lot easier. Google, therefore, has a number of mechanisms (some patented) that cause all hell to break loose if you make the kind of changes to your site that you'd use to incrementally improve it's SEO.

It's one of the reason we are stuck with crappy sites like answers.com, w3schools, and wrongdiagnosis, because once a site like that is successful, the operators are loathe to make any changes lest their rankings drop.

Re: Matt Cutts is looking for scraper sites

#82
post #77

Earlier quoted context omitted.

One man's "scrapper" is another man's "aggregator". How do you think Google would view my site if I wrapped Wikipedia's content, with back link and ran my own ads alongside that content? I would imagine not very positively. Also, is it okay that a bigger entity scrapes my content just because they send me traffic? You might not want to bite the hand that feeds you, but it still doesn't make it right.

Google does not reproduce whole articles, only short excerpts to help searchers decide whether it's relevant to what they're looking for - and with clear indication of the source and in a context where it's understood that Google is showing the blurb only to pointing to the source where it was found. This is technically scraping but it's hardly comparable to the bottom-feeders that plagiarize for money. (Edit: accord…

You're comparing what Google does to another extreme when you say things like "bottom-feeders that plagiarize for money"

Surely you don't believe that all "scapers" are bottom feeders? It's like saying every criminal is a murderer. There's a whole bunch of grey area in between, and this is where the criticism of Google's harsh penalties is valid.

Re: Matt Cutts is looking for scraper sites

#83

It's a funny quip but it's getting more attention than the important piece of news it highlights. Google is finally doing something about scraping sites doing better in search results than original creators. Good. Many people don't write for money, to put ads on their website, or as part of some "content marketing" campaign. All they want is a little recognition. A boost in positioning on the SERP means we will be ge…

It is highlighting the more important aspect - Google is the largest scraping site in the world.

Re: Matt Cutts is looking for scraper sites

#84
post #76

Earlier quoted context omitted.

I'm not sure which point you're trying to make. Did you look at op's submission?

I assumed you thought scraping wikipedia and putting it on top of the search results was unethical. (The other alternative - that punishing scraper sites is unethical - seemed unlikely). So the fact that you prefer DDG because of this seemed weird, considering DDG does the same thing.

DDG is clearly attributing to Wikipedia and is not cloaking the link behind a redirect.

Re: Matt Cutts is looking for scraper sites

#85
post #72

Earlier quoted context omitted.

That's one way of looking at it, on the other hand, they link to the original URL, passing traffic back to the original source. Most "scraper" sites take the content, wrap it in their own similar outer layer, and try to take ad revenue. E.g. I've seen my own StackOverflow answers copied, word for word, to a scraper site and presented under a made-up name.

They don't actually link to the wikipedia URL. They mask a link that leads to another Google page "/url?sa=t&rct=j&q=&...." which in turn responds with a 200 OK page that redirects to Wikipedia. Sure it passes the keywords etc. But this likely reduces the number of people visiting Wikipedia, while increasing Google's ad revenues, if anyone but Google did this they'd be potential blacklisted by Google.

Actually, they do link to the wikipedia URL.

href="http://en.wikipedia.org/wiki/Scraper_site" appears directly in the source code of that web page.

It also has an onmousedown handler that rewrites the URL to point at Google, so they can tell which link you clicked, to improve their ranking system. And Google works very closely with sites to make sure the sites know how to understand the referrals.

Re: Matt Cutts is looking for scraper sites

#86

Earlier quoted context omitted.

"Google is finally doing something about scraping" I hope this is genuine and not a disingenuous diversion on Google's part. The fact that the Huffington Post still ranks very high for trendy searches makes me wonder. As usual, follow the money: the scraping sites exist to make money, often through Google's advertising; Google gets a cut. The original content is often on sites with no advertising or real traffic, fro…

Huffington Post isn't a scraper site. Aside from the original content they produce, they republish blog posts with permission from the authors. If you have an example of Huffington Post literally cut-and-pasting content from someone without attribution, please share. I also assume that by "HuffPo investors" you mean AOL? Huffington Post is a fully owned subsidiary. (Disclosure: I consult for Huffington Post)

You are right. Please see my second edit in my comment.

Re: Matt Cutts is looking for scraper sites

#87

It's a funny quip but it's getting more attention than the important piece of news it highlights. Google is finally doing something about scraping sites doing better in search results than original creators. Good. Many people don't write for money, to put ads on their website, or as part of some "content marketing" campaign. All they want is a little recognition. A boost in positioning on the SERP means we will be ge…

"Google is finally doing something about scraping" I hope this is genuine and not a disingenuous diversion on Google's part. The fact that the Huffington Post still ranks very high for trendy searches makes me wonder. As usual, follow the money: the scraping sites exist to make money, often through Google's advertising; Google gets a cut. The original content is often on sites with no advertising or real traffic, fro…

I covered this a few months back while it all went down with the Verge and HuffPo and how our social search engine algorithm accounted for this while Google did not.

http://theenginuity.com/blog/how-a-copied-excerpt-of-a-story...

Re: Matt Cutts is looking for scraper sites

#88
Hey guys you know this is meant to be humorous right? I honestly can't believe that people here are saying Google is a scraper site and complaining about "hypocrisy". No more caching! When I search Google I want them to freshly crawl the web and get back to me in a day or two with my results.

Re: Matt Cutts is looking for scraper sites

#89

It's a funny quip but it's getting more attention than the important piece of news it highlights. Google is finally doing something about scraping sites doing better in search results than original creators. Good. Many people don't write for money, to put ads on their website, or as part of some "content marketing" campaign. All they want is a little recognition. A boost in positioning on the SERP means we will be ge…

"Google is finally doing something about scraping" I hope this is genuine and not a disingenuous diversion on Google's part. The fact that the Huffington Post still ranks very high for trendy searches makes me wonder. As usual, follow the money: the scraping sites exist to make money, often through Google's advertising; Google gets a cut. The original content is often on sites with no advertising or real traffic, fro…

HuffPo definitely has original reporters, my friend is one of them:

http://www.huffingtonpost.com/betsy-isaacson/

She writes good articles for the general public about tech in general and things like net neutrality and aaron swartz

Re: Matt Cutts is looking for scraper sites

#90

Earlier quoted context omitted.

"Google is finally doing something about scraping" I hope this is genuine and not a disingenuous diversion on Google's part. The fact that the Huffington Post still ranks very high for trendy searches makes me wonder. As usual, follow the money: the scraping sites exist to make money, often through Google's advertising; Google gets a cut. The original content is often on sites with no advertising or real traffic, fro…

I covered this a few months back while it all went down with the Verge and HuffPo and how our social search engine algorithm accounted for this while Google did not. http://theenginuity.com/blog/how-a-copied-excerpt-of-a-story...

You are correct, and the parent post before you is also correct.

Google's algorithm put a great deal of value on domain names, which provided a strong incentive for owners of a strong domain name to pump out low quality content. Low quality content can be the example you give, it can be "re-authoring" someone elses article, or is can be blatant copy and paste which is generally avoided due to the obviousness.

When a media property, such as the Huffington Post, pumps out this volume of low quality content the advertising revenue can subsidize the cost of paying for original journalist content.

On another note, take a look at The Daily Mail. They pump out timely news pieces so quickly that they are covered in typos and sometimes can't even keep left and right straight in photo captions.

Post reply on HN