Live data from Hacker News

Matt Cutts is looking for scraper sites

twitter.com

91–100 of 120 posts

Re: Matt Cutts is looking for scraper sites

#92
post #77

Earlier quoted context omitted.

One man's "scrapper" is another man's "aggregator". How do you think Google would view my site if I wrapped Wikipedia's content, with back link and ran my own ads alongside that content? I would imagine not very positively. Also, is it okay that a bigger entity scrapes my content just because they send me traffic? You might not want to bite the hand that feeds you, but it still doesn't make it right.

Google does not reproduce whole articles, only short excerpts to help searchers decide whether it's relevant to what they're looking for - and with clear indication of the source and in a context where it's understood that Google is showing the blurb only to pointing to the source where it was found. This is technically scraping but it's hardly comparable to the bottom-feeders that plagiarize for money. (Edit: accord…

You are at least partially incorrect.

Last year Google was testing reproducing entire Wikipedia articles within their site for their mobile site. You could read the full article within going to Wikipedia (allowed by Creative Commons, of course.) Between that and what they did with Google Images, I would say this reveals intention and is the direction web publishers should expect Google to be headed in.

In order for Google to continue to meet their growth targets they must increase the percentage of outgoing click from free to paid.

Re: Matt Cutts is looking for scraper sites

#94

It's a funny quip but it's getting more attention than the important piece of news it highlights. Google is finally doing something about scraping sites doing better in search results than original creators. Good. Many people don't write for money, to put ads on their website, or as part of some "content marketing" campaign. All they want is a little recognition. A boost in positioning on the SERP means we will be ge…

Scrapping is only possible because the value of original content creation is 0% and the value of the toll collecting aggregation is 100% via ads. This is a structural problem in the universe Goggle created, not some anomaly that can be handled by ranking.

Re: Matt Cutts is looking for scraper sites

#95

It's a funny quip but it's getting more attention than the important piece of news it highlights. Google is finally doing something about scraping sites doing better in search results than original creators. Good. Many people don't write for money, to put ads on their website, or as part of some "content marketing" campaign. All they want is a little recognition. A boost in positioning on the SERP means we will be ge…

I've always wondered about news organizations. The vast majority of them have websites full of regurgitated content. They don't usually discover and report their own news, they feed off Associated Press, Reuters and often times individuals posting videos on youtube or articles in their blogs. It would be nice to see them not dominate news on the internet by simply engaging in re-writing what someone else originated.

Is it content curation? Don't know. Everyone is reporting pretty much exactly the same things. They can't all be the original source. Who gets the juice?

Re: Matt Cutts is looking for scraper sites

#96
Google is most certainly crossing the line here.

1. They are not only doing this with wikipedia, but with many, many sites: "what is the smallest cell in the human body", "what is the biggest planet in the solar system".

2. The sites they chose to link are not always the highest quality sites, such as the two examples above- why are these websites being featured?

3. Many times, the user will get their answer right then and there, and be done with the search process. The site misses a visitor. In spite of these type of questions being "facts", someone took the time to organize and give context to these "facts". Turning facts into useful, consumable, content costs money. Google should not be taking visitors away from these sites.

4. There should be public information on the CTR of these snippets. See if it helps or hurts the user.

5. Google is abusing its power as a major search engine to reinforce structuring rules, such as microformats. With these rules, webmasters are giving more and more semantic meaning to their content, which means Google has an easier time completing their knowledge graph. They might link to the source site for a while, but there is no good argument for linking back to wikipedia to attribute the fact that Jupiter is the largest planet, since it's a fact, just like 2+2 is 4 (no attribution).

6. Google is all about ML/NLP/AI driven knowledge. But in reality they are turning all of the internet content creators into a giant sweat shop for their knowledge graph. This is not fair, and sooner or later it will come back to bite them.

Re: Matt Cutts is looking for scraper sites

#97

It's a funny quip but it's getting more attention than the important piece of news it highlights. Google is finally doing something about scraping sites doing better in search results than original creators. Good. Many people don't write for money, to put ads on their website, or as part of some "content marketing" campaign. All they want is a little recognition. A boost in positioning on the SERP means we will be ge…

Yes, because only Google can earn advertising money by scrapping sites.

Re: Matt Cutts is looking for scraper sites

#98
post #77

Earlier quoted context omitted.

One man's "scrapper" is another man's "aggregator". How do you think Google would view my site if I wrapped Wikipedia's content, with back link and ran my own ads alongside that content? I would imagine not very positively. Also, is it okay that a bigger entity scrapes my content just because they send me traffic? You might not want to bite the hand that feeds you, but it still doesn't make it right.

Google does not reproduce whole articles, only short excerpts to help searchers decide whether it's relevant to what they're looking for - and with clear indication of the source and in a context where it's understood that Google is showing the blurb only to pointing to the source where it was found. This is technically scraping but it's hardly comparable to the bottom-feeders that plagiarize for money. (Edit: accord…

Google does not reproduce whole articles, only short excerpts to help searchers decide whether it's relevant to what they're looking for

And what about, for example, Google's image search tool, where the image itself might be what their user is searching for, and where Google controversially changed their system a little while ago to show full-size images in-SERP and de-emphasize forwarding search users to the original source? Or Google Cache, if it's reproducing material that has since been taken down deliberately from the original source?

To add insult to injury, some Google services still appear to rely on the original source's bandwidth to serve things like images (not to mention avoiding a certain legal argument about copyright infringement), thus violating the basic principle of netiquette that has been good manners ever since people actually used the word netiquette that you don't hotlink other people's stuff on your site.

Re: Matt Cutts is looking for scraper sites

#99
post #43

Earlier quoted context omitted.

Do the terms of the license allow for this kind of abuse? Just because something is CC doesn't mean you can do whatever you want with it.

Yes, they do; it's not abuse when you're given explicit permission. CC BY-SA means you can do whatever you want with it as long as you attribute the source as specified.

CC BY-SA is short for Creative Commons Attribution Share-Alike. BY means you must attribute, and SA means you must license any distributed derivative works under the same license (copyleft). Attribution on its own is not enough.
Post reply on HN