Live data from Hacker News

Matt Cutts is looking for scraper sites

twitter.com

31–40 of 120 posts

Re: Matt Cutts is looking for scraper sites

#31
post #27

Wikipedia's database is public and used by Google with permission. You can probably use it for your projects, too. So this is neither scraping, nor against the rules. Here are dumps in SQL and XML format: http://dumps.wikimedia.org/enwiki/ Ps- Yes the original post was meant to be funny and it was; I do have a sense of humor. :)

He's talking about outranking the true original source of the content in search results. You most certainly cannot create your own site that consists only of excerpts from Wikipedia, if you wish to remain on Google's search results. Copyrights are irrelevant to this.

What's bad though, is that Google isn't just lowering the rankings of non-original content pages now (including any kind of legitimate curation sites.) They're marking the entire domains of new curation sites as "pure spam" and de-listing them from Google entirely, and punishing anyone who's linked to them.

This is having the effect of sending a clear message to developers -- stay far away from Google's territory of recommending third party content to people, no matter how you do it.

Re: Matt Cutts is looking for scraper sites

#32
Who is deemed to be the scraper? The site that get crawled and indexed first, and ranks better, or the site that ranks well but has sites with scraped content that doesn't rank as well?

Matt is looking for scapers that rank better than the original, basically meaning that they have higher PageRank and more links.

Re: Matt Cutts is looking for scraper sites

#33

Scrapers lift the full content, wholesale, without attribution. You may as will just show http://images.google.com and complain that it's scraping. Or http://news.google.com . In general, do you think Wikipedia gets more traffic because Google exists, or do you think Google gets more traffic because Wikipedia exists? Meaning, which affect is larger? I'm pretty sure the answer to this is obvious. And if more scrapers…

That's one way of looking at it, on the other hand, they link to the original URL, passing traffic back to the original source. Most "scraper" sites take the content, wrap it in their own similar outer layer, and try to take ad revenue. E.g. I've seen my own StackOverflow answers copied, word for word, to a scraper site and presented under a made-up name.

By having a tl;dr about the actual Wikipedia page, there is no need for the user to click on the link. Following what you're saying, Google as wrapped it in their own layer, and trying to take ad revenue.

Re: Matt Cutts is looking for scraper sites

#34
post #6

I wonder how Google chooses which Wikipedia articles they scrape and which ones they don't. In testing, they definitely don't seem to scrape every article: http://i.imgur.com/ujDqZhB.png

This is a good question...I've long since surmised that Google has a set of heuristics for every site that has an API that allows for easy domain-specific ranking. With Wikipedia, you have number of page edits, frequency of page edits, and (to an extent) quality of recent page edits. StackOverflow provides an even easier metric for what's considered high quality, and Google appears to apply its own layer on top of that (and in my non-scientific perception, looking something up by Google is almost always more fruitful on the first search than by going directly to SO)

Re: Matt Cutts is looking for scraper sites

#35
post #19
post #16

Cue Bing, DuckDuckGo and any other search engine (except Google, of course) being Google-killed for "scraping". It's the perfect plan!

Google recently flagged my content-curation startup as "Pure Spam", even though it only takes small snippets from the original sources, is 100% human curated, and always links back to the true original source. Not only are the curated pages blocked, but the entire domain is blocked as "pure spam". People who use Google to find a domain instead of typing the full URL now can't find it anywhere. These assholes are just…

Curious - Which site is that?

Re: Matt Cutts is looking for scraper sites

#37
post #29
post #15

There are lots of places where Google decides to "help" me, but sometimes I just want search results. Other times, I actually like getting the curated content (e.g. search for "delta 3810"). Is there a way to disable this? EDIT: I should also note that I'm one of those who switched over to DuckDuckGo for privacy reasons, so I don't see these results as often now.

The content they do provide is often so bad it is almost embarrassing. Search for "Russia" and you get a completely useless map and a list of random facts. It may be useful to a child researching geography but for me it is just annoying. I want content that is curated by people who actually understand the subject. I would pay for a search engine designed by someone who understands my industry. The Google algorithm on…

I'm experimenting with paying someone to do a bit of research for me.

To give some idea I've asked for a list of URLs to documents covering best current practice for suicide prevention in Gloucestershire and Herefordshire; to include national level NHS and NICE guidance, DoH guidance, anything from Gloucestershire and Herefordshire, and anything recognised as excellent from anywhere else in the country. If possible I also want a list of protocols used in schools, care homes, etc.

It's probably something you could risk on MTurk. Perhaps Bountify.com could expand to this kind of simple websearching.

Obvious drawbacks include delay between starting the search and getting the results, and cost, and having to trust some random person to not miss stuff.

I don't know if there's anything similar to "clippings services" either where you'd provide them with list of types of stories you'd want, and they'd read all the newspapers and clip any relevant stories and post them to you.

Re: Matt Cutts is looking for scraper sites

#38

Scrapers lift the full content, wholesale, without attribution. You may as will just show http://images.google.com and complain that it's scraping. Or http://news.google.com . In general, do you think Wikipedia gets more traffic because Google exists, or do you think Google gets more traffic because Wikipedia exists? Meaning, which affect is larger? I'm pretty sure the answer to this is obvious. And if more scrapers…

One man's "scrapper" is another man's "aggregator".

How do you think Google would view my site if I wrapped Wikipedia's content, with back link and ran my own ads alongside that content? I would imagine not very positively.

Also, is it okay that a bigger entity scrapes my content just because they send me traffic? You might not want to bite the hand that feeds you, but it still doesn't make it right.

Re: Matt Cutts is looking for scraper sites

#39
It's a funny quip but it's getting more attention than the important piece of news it highlights. Google is finally doing something about scraping sites doing better in search results than original creators. Good.

Many people don't write for money, to put ads on their website, or as part of some "content marketing" campaign. All they want is a little recognition. A boost in positioning on the SERP means we will be getting useful stuff at no cost.

And there are genuine replies there. Ryan Jones[1] even got the scrapers to confess their sins[2].

[1] https://twitter.com/RyanJones/status/439123533349015553

[2] https://www.google.com/search?q=%20%22istwfn%22+%22stole+thi...

Post reply on HN