Scrapers lift the full content, wholesale, without attribution. You may as will just show http://images.google.com and complain that it's scraping. Or http://news.google.com . In general, do you think Wikipedia gets more traffic because Google exists, or do you think Google gets more traffic because Wikipedia exists? Meaning, which affect is larger? I'm pretty sure the answer to this is obvious. And if more scrapers…
That's one way of looking at it, on the other hand, they link to the original URL, passing traffic back to the original source. Most "scraper" sites take the content, wrap it in their own similar outer layer, and try to take ad revenue. E.g. I've seen my own StackOverflow answers copied, word for word, to a scraper site and presented under a made-up name.
Matt Cutts is looking for scraper sites
41–50 of 120 posts
Re: Matt Cutts is looking for scraper sites
#42Bah, what would I possibly need with a scraped definition that 1) Hasn't been chunked into 20 pieces of varying grammatical structure which are automatically matched to corresponding questions 2) Hasn't been subsequently pasted over a slideshow of completely irrelevant stock photos in bold, white font 3) Isn't accompanied by a grid of ~30 vaguely related questions helpfully linked to similar pages and tastefully deco…
Re: Matt Cutts is looking for scraper sites
#43Earlier quoted context omitted.
That's one way of looking at it, on the other hand, they link to the original URL, passing traffic back to the original source. Most "scraper" sites take the content, wrap it in their own similar outer layer, and try to take ad revenue. E.g. I've seen my own StackOverflow answers copied, word for word, to a scraper site and presented under a made-up name.
StackOverflow actually allows this; all their data is Creative Commons licensed, and they publish the full database dump on the Internet Archive. https://archive.org/details/stackexchange
Just because something is CC doesn't mean you can do whatever you want with it.
Re: Matt Cutts is looking for scraper sites
#44Easy. Search for a programming related question. After the result from stackoverflow you'll find dozens of scraper sites.
Re: Matt Cutts is looking for scraper sites
#45Wikipedia's database is public and used by Google with permission. You can probably use it for your projects, too. So this is neither scraping, nor against the rules. Here are dumps in SQL and XML format: http://dumps.wikimedia.org/enwiki/ Ps- Yes the original post was meant to be funny and it was; I do have a sense of humor. :)
He's talking about outranking the true original source of the content in search results. You most certainly cannot create your own site that consists only of excerpts from Wikipedia, if you wish to remain on Google's search results. Copyrights are irrelevant to this. What's bad though, is that Google isn't just lowering the rankings of non-original content pages now (including any kind of legitimate curation sites.)…
Re: Matt Cutts is looking for scraper sites
#46Earlier quoted context omitted.
Google recently flagged my content-curation startup as "Pure Spam", even though it only takes small snippets from the original sources, is 100% human curated, and always links back to the true original source. Not only are the curated pages blocked, but the entire domain is blocked as "pure spam". People who use Google to find a domain instead of typing the full URL now can't find it anywhere. These assholes are just…
Curious - Which site is that?
This reminds me. Google is really killing the "release early and release often" approach, if people will now have to do a ton of SEO learning and tweaking to avoid having your MVP permanently banned at launch day.
Re: Matt Cutts is looking for scraper sites
#47Re: Matt Cutts is looking for scraper sites
#48Earlier quoted context omitted.
That's one way of looking at it, on the other hand, they link to the original URL, passing traffic back to the original source. Most "scraper" sites take the content, wrap it in their own similar outer layer, and try to take ad revenue. E.g. I've seen my own StackOverflow answers copied, word for word, to a scraper site and presented under a made-up name.
By having a tl;dr about the actual Wikipedia page, there is no need for the user to click on the link. Following what you're saying, Google as wrapped it in their own layer, and trying to take ad revenue.
Re: Matt Cutts is looking for scraper sites
#49There are lots of places where Google decides to "help" me, but sometimes I just want search results. Other times, I actually like getting the curated content (e.g. search for "delta 3810"). Is there a way to disable this? EDIT: I should also note that I'm one of those who switched over to DuckDuckGo for privacy reasons, so I don't see these results as often now.
The content they do provide is often so bad it is almost embarrassing. Search for "Russia" and you get a completely useless map and a list of random facts. It may be useful to a child researching geography but for me it is just annoying. I want content that is curated by people who actually understand the subject. I would pay for a search engine designed by someone who understands my industry. The Google algorithm on…
"I would pay for a search engine designed by someone who understands my industry"
What industry is that? And how would Google guess or know your industry unless you tell them?
Re: Matt Cutts is looking for scraper sites
#50Earlier quoted context omitted.
StackOverflow actually allows this; all their data is Creative Commons licensed, and they publish the full database dump on the Internet Archive. https://archive.org/details/stackexchange
Do the terms of the license allow for this kind of abuse? Just because something is CC doesn't mean you can do whatever you want with it.