Live data from Hacker News

Matt Cutts is looking for scraper sites

twitter.com

11–20 of 120 posts

Re: Matt Cutts is looking for scraper sites

#11
Scrapers lift the full content, wholesale, without attribution.

You may as will just show http://images.google.com and complain that it's scraping. Or http://news.google.com.

In general, do you think Wikipedia gets more traffic because Google exists, or do you think Google gets more traffic because Wikipedia exists? Meaning, which affect is larger? I'm pretty sure the answer to this is obvious.

And if more scrapers donated millions to the site they scrape from, the world would be a much better place.

http://wikimediafoundation.org/wiki/Press_releases/Wikimedia...

Re: Matt Cutts is looking for scraper sites

#12
post #9

Easy. Search for a programming related question. After the result from stackoverflow you'll find dozens of scraper sites.

If you read correctly, the task is to report scraper sites that rank HIGHER than the original site. Not the case in your scenario.

Re: Matt Cutts is looking for scraper sites

#15
There are lots of places where Google decides to "help" me, but sometimes I just want search results. Other times, I actually like getting the curated content (e.g. search for "delta 3810"). Is there a way to disable this?

EDIT: I should also note that I'm one of those who switched over to DuckDuckGo for privacy reasons, so I don't see these results as often now.

Re: Matt Cutts is looking for scraper sites

#17

Scrapers lift the full content, wholesale, without attribution. You may as will just show http://images.google.com and complain that it's scraping. Or http://news.google.com . In general, do you think Wikipedia gets more traffic because Google exists, or do you think Google gets more traffic because Wikipedia exists? Meaning, which affect is larger? I'm pretty sure the answer to this is obvious. And if more scrapers…

If Google only needs to visit Wikipedia's "scraper" page once a day or less, but serves it out to others with attribution, isn't that helping Wikipedia by lowering traffic COSTS?

Re: Matt Cutts is looking for scraper sites

#19
post #16

Cue Bing, DuckDuckGo and any other search engine (except Google, of course) being Google-killed for "scraping". It's the perfect plan!

Google recently flagged my content-curation startup as "Pure Spam", even though it only takes small snippets from the original sources, is 100% human curated, and always links back to the true original source.

Not only are the curated pages blocked, but the entire domain is blocked as "pure spam". People who use Google to find a domain instead of typing the full URL now can't find it anywhere.

These assholes are just being anti-competitive now.

Re: Matt Cutts is looking for scraper sites

#20
post #17

Scrapers lift the full content, wholesale, without attribution. You may as will just show http://images.google.com and complain that it's scraping. Or http://news.google.com . In general, do you think Wikipedia gets more traffic because Google exists, or do you think Google gets more traffic because Wikipedia exists? Meaning, which affect is larger? I'm pretty sure the answer to this is obvious. And if more scrapers…

If Google only needs to visit Wikipedia's "scraper" page once a day or less, but serves it out to others with attribution, isn't that helping Wikipedia by lowering traffic COSTS?

BUT it gives full attribution to http://en.wikipedia.org/wiki/Scraper_site???
Post reply on HN