Live data from Hacker News

Matt Cutts is looking for scraper sites

twitter.com

71–80 of 120 posts

Re: Matt Cutts is looking for scraper sites

#71

Earlier quoted context omitted.

That's one way of looking at it, on the other hand, they link to the original URL, passing traffic back to the original source. Most "scraper" sites take the content, wrap it in their own similar outer layer, and try to take ad revenue. E.g. I've seen my own StackOverflow answers copied, word for word, to a scraper site and presented under a made-up name.

StackOverflow actually allows this; all their data is Creative Commons licensed, and they publish the full database dump on the Internet Archive. https://archive.org/details/stackexchange

Interesting, from the file sizes you can quickly gauge the relative popularity of each subject.

Re: Matt Cutts is looking for scraper sites

#72

Scrapers lift the full content, wholesale, without attribution. You may as will just show http://images.google.com and complain that it's scraping. Or http://news.google.com . In general, do you think Wikipedia gets more traffic because Google exists, or do you think Google gets more traffic because Wikipedia exists? Meaning, which affect is larger? I'm pretty sure the answer to this is obvious. And if more scrapers…

That's one way of looking at it, on the other hand, they link to the original URL, passing traffic back to the original source. Most "scraper" sites take the content, wrap it in their own similar outer layer, and try to take ad revenue. E.g. I've seen my own StackOverflow answers copied, word for word, to a scraper site and presented under a made-up name.

They don't actually link to the wikipedia URL. They mask a link that leads to another Google page "/url?sa=t&rct=j&q=&...." which in turn responds with a 200 OK page that redirects to Wikipedia.

Sure it passes the keywords etc. But this likely reduces the number of people visiting Wikipedia, while increasing Google's ad revenues, if anyone but Google did this they'd be potential blacklisted by Google.

Re: Matt Cutts is looking for scraper sites

#73

It's a funny quip but it's getting more attention than the important piece of news it highlights. Google is finally doing something about scraping sites doing better in search results than original creators. Good. Many people don't write for money, to put ads on their website, or as part of some "content marketing" campaign. All they want is a little recognition. A boost in positioning on the SERP means we will be ge…

"Google is finally doing something about scraping" I hope this is genuine and not a disingenuous diversion on Google's part. The fact that the Huffington Post still ranks very high for trendy searches makes me wonder. As usual, follow the money: the scraping sites exist to make money, often through Google's advertising; Google gets a cut. The original content is often on sites with no advertising or real traffic, fro…

Is their copy and paste verbatim?

Re: Matt Cutts is looking for scraper sites

#74

It's a funny quip but it's getting more attention than the important piece of news it highlights. Google is finally doing something about scraping sites doing better in search results than original creators. Good. Many people don't write for money, to put ads on their website, or as part of some "content marketing" campaign. All they want is a little recognition. A boost in positioning on the SERP means we will be ge…

"Google is finally doing something about scraping" I hope this is genuine and not a disingenuous diversion on Google's part. The fact that the Huffington Post still ranks very high for trendy searches makes me wonder. As usual, follow the money: the scraping sites exist to make money, often through Google's advertising; Google gets a cut. The original content is often on sites with no advertising or real traffic, fro…

That would be a pretty blatant copyright violation. Can you provide an example to substantiate the claim?

Re: Matt Cutts is looking for scraper sites

#75

It's a funny quip but it's getting more attention than the important piece of news it highlights. Google is finally doing something about scraping sites doing better in search results than original creators. Good. Many people don't write for money, to put ads on their website, or as part of some "content marketing" campaign. All they want is a little recognition. A boost in positioning on the SERP means we will be ge…

"Google is finally doing something about scraping" I hope this is genuine and not a disingenuous diversion on Google's part. The fact that the Huffington Post still ranks very high for trendy searches makes me wonder. As usual, follow the money: the scraping sites exist to make money, often through Google's advertising; Google gets a cut. The original content is often on sites with no advertising or real traffic, fro…

Huffington Post isn't a scraper site. Aside from the original content they produce, they republish blog posts with permission from the authors. If you have an example of Huffington Post literally cut-and-pasting content from someone without attribution, please share.

I also assume that by "HuffPo investors" you mean AOL? Huffington Post is a fully owned subsidiary.

(Disclosure: I consult for Huffington Post)

Re: Matt Cutts is looking for scraper sites

#76
post #52

Earlier quoted context omitted.

https://duckduckgo.com/?q=scraper+site

I'm not sure which point you're trying to make. Did you look at op's submission?

I assumed you thought scraping wikipedia and putting it on top of the search results was unethical. (The other alternative - that punishing scraper sites is unethical - seemed unlikely). So the fact that you prefer DDG because of this seemed weird, considering DDG does the same thing.

Re: Matt Cutts is looking for scraper sites

#77

Scrapers lift the full content, wholesale, without attribution. You may as will just show http://images.google.com and complain that it's scraping. Or http://news.google.com . In general, do you think Wikipedia gets more traffic because Google exists, or do you think Google gets more traffic because Wikipedia exists? Meaning, which affect is larger? I'm pretty sure the answer to this is obvious. And if more scrapers…

One man's "scrapper" is another man's "aggregator". How do you think Google would view my site if I wrapped Wikipedia's content, with back link and ran my own ads alongside that content? I would imagine not very positively. Also, is it okay that a bigger entity scrapes my content just because they send me traffic? You might not want to bite the hand that feeds you, but it still doesn't make it right.

Google does not reproduce whole articles, only short excerpts to help searchers decide whether it's relevant to what they're looking for - and with clear indication of the source and in a context where it's understood that Google is showing the blurb only to pointing to the source where it was found.

This is technically scraping but it's hardly comparable to the bottom-feeders that plagiarize for money. (Edit: according to 'pud' on this page, Google uses a Wikipedia index so it's not scraping, but it is in the case of other sites that Google indexes.)

And yes, it's OK both legally and ethically if you do the same to Wikipedia - like Google that is, just for indexing purposes and not using whole articles.

Re: Matt Cutts is looking for scraper sites

#78
post #43

Earlier quoted context omitted.

StackOverflow actually allows this; all their data is Creative Commons licensed, and they publish the full database dump on the Internet Archive. https://archive.org/details/stackexchange

Do the terms of the license allow for this kind of abuse? Just because something is CC doesn't mean you can do whatever you want with it.

No, attribution is required.

Re: Matt Cutts is looking for scraper sites

#79

It's a funny quip but it's getting more attention than the important piece of news it highlights. Google is finally doing something about scraping sites doing better in search results than original creators. Good. Many people don't write for money, to put ads on their website, or as part of some "content marketing" campaign. All they want is a little recognition. A boost in positioning on the SERP means we will be ge…

"Google is finally doing something about scraping" I hope this is genuine and not a disingenuous diversion on Google's part. The fact that the Huffington Post still ranks very high for trendy searches makes me wonder. As usual, follow the money: the scraping sites exist to make money, often through Google's advertising; Google gets a cut. The original content is often on sites with no advertising or real traffic, fro…

There's a big difference between writing an article based on another article you read, and web scrapers.

Not that I particularly feel like defending the Huffington Post, but they're not a web scraper.

Re: Matt Cutts is looking for scraper sites

#80

It's a funny quip but it's getting more attention than the important piece of news it highlights. Google is finally doing something about scraping sites doing better in search results than original creators. Good. Many people don't write for money, to put ads on their website, or as part of some "content marketing" campaign. All they want is a little recognition. A boost in positioning on the SERP means we will be ge…

"Google is finally doing something about scraping" I hope this is genuine and not a disingenuous diversion on Google's part. The fact that the Huffington Post still ranks very high for trendy searches makes me wonder. As usual, follow the money: the scraping sites exist to make money, often through Google's advertising; Google gets a cut. The original content is often on sites with no advertising or real traffic, fro…

Google-search for any hot topic in the news, say the name of some misbehaving pop star. See the HuffPo result near the top of the page. Look down to see several results from real newspapers.

Many newspapers get a lot of their content from syndication services like Reuters. You may be seeing similar content because lazy editorial assistants just copied out a reuters story verbatim, slapped a pic on it and put it up at multiple organisations, not because HuffPo is scraping other sites. Do you have an example of this sort of thing you can point to? It'd be interesting to trace the origin of the content.

Post reply on HN