Live data from Hacker News

Matt Cutts is looking for scraper sites

twitter.com

111–120 of 120 posts

Re: Matt Cutts is looking for scraper sites

#111
post #53

Google is taking Cognitive Dissonance to a new level: It's okay for them to scrape every single site, download its content and images, and cache it on their servers, and run their AD platform on top of it. but that's not enough, they would still like to impose their rules and punish people who do the same thing.

"Cognitive dissonance" doesn't mean the same thing as "hypocrisy."

Re: Matt Cutts is looking for scraper sites

#113
post #45
post #31

Earlier quoted context omitted.

He's talking about outranking the true original source of the content in search results. You most certainly cannot create your own site that consists only of excerpts from Wikipedia, if you wish to remain on Google's search results. Copyrights are irrelevant to this. What's bad though, is that Google isn't just lowering the rankings of non-original content pages now (including any kind of legitimate curation sites.)…

Could you show an example of such a legitimate new curation site please?

Not new but a legitimate curation site - http://hypem.com/

Re: Matt Cutts is looking for scraper sites

#114
post #62

Bah, what would I possibly need with a scraped definition that 1) Hasn't been chunked into 20 pieces of varying grammatical structure which are automatically matched to corresponding questions 2) Hasn't been subsequently pasted over a slideshow of completely irrelevant stock photos in bold, white font 3) Isn't accompanied by a grid of ~30 vaguely related questions helpfully linked to similar pages and tastefully deco…

If you use Chrome, there is a "Personal Blacklist" extension that does essentially what the manual blacklist used to do.

"If you use Chrome" is Internet Explorer ActiveX controls redux

Re: Matt Cutts is looking for scraper sites

#115
post #108

Earlier quoted context omitted.

It seems a bit odd for you to be so cagey about the identity of this "original site" while at the same time lamenting that they aren't getting the traffic they deserve. Why don't you tell us who they are?

It's because I don't speak for the owners of the site and I'd rather make sure they don't mind me putting it out there like this. I could let you know privately, if you'd like to check my story for yourself, though.

Why on earth would they mind?

Re: Matt Cutts is looking for scraper sites

#116
post #115

Earlier quoted context omitted.

It's because I don't speak for the owners of the site and I'd rather make sure they don't mind me putting it out there like this. I could let you know privately, if you'd like to check my story for yourself, though.

Why on earth would they mind?

I'm not sure if they do mind. I do know that their relationship with Google is important to them when it comes to their much larger and more successful projects and that this site has been mostly left behind, so they may not want to bring it up in the context of this Hacker News post, even in the unlikely case that it resulted in this site getting its traffic back. Why not just email me and I'll show you a simple content site with minimal traffic, not using any black or gray-hat SEO tactics, with high-quality, original (to the Web) content, referenced in thousands of Wikipedia articles and you can decide for yourself if my post was truthful.

Re: Matt Cutts is looking for scraper sites

#118
I do not understand the Wikipedia definition of "scraper site".

By this definition webcache.googleusercontent.com qualifies.

It is a full copy of every site GoogleBot scrapes.

Google gives attrition to the original source, but if this isn't "scraping", what is?

They have been sued for this, and they've won. The benefits of a decent search engine outweigh the burden of infringing the copyrights of others. At least where Google and other search engines that cache websites are concerned.

Re: Matt Cutts is looking for scraper sites

#119
post #118

I do not understand the Wikipedia definition of "scraper site". By this definition webcache.googleusercontent.com qualifies. It is a full copy of every site GoogleBot scrapes. Google gives attrition to the original source, but if this isn't "scraping", what is? They have been sued for this, and they've won. The benefits of a decent search engine outweigh the burden of infringing the copyrights of others. At least whe…

s/attrition/attribution/

Re: Matt Cutts is looking for scraper sites

#120
post #115

Earlier quoted context omitted.

Why on earth would they mind?

I'm not sure if they do mind. I do know that their relationship with Google is important to them when it comes to their much larger and more successful projects and that this site has been mostly left behind, so they may not want to bring it up in the context of this Hacker News post, even in the unlikely case that it resulted in this site getting its traffic back. Why not just email me and I'll show you a simple con…

Could you get permission from the owners and publish it here?
Post reply on HN