Live data from Hacker News

Google doesn't recognise or penalise stolen content

pi-datametrics.com

61–70 of 80 posts

Re: Google doesn't recognise or penalise stolen content

#61
What a rubbish article.. It doesn't fly for a second under copyright law. Google is entirely within their rights doing what they're doing. The onus isn't on Google to detect the infringing content.

For anyone interested in copyright and legal issues, I'd recommend checking out techdirt.com. They have a great starter section at https://www.techdirt.com/blog/?tag=techdirt+feature, and they cover legal, copyright, patent, surveillance and all sorts of related topics. High quality journalism.

Re: Google doesn't recognise or penalise stolen content

#62

I think this is pretty fair on Google's part. How could you possibly figure out who owned content? What if I published a book, it was copy-pasted in blogs, and then later I put it somewhere crawlable by Google? You certainly can't just say "first time we saw it, that's the proper owner". It would either require a massive amount of manual QA to get right (and even then, there are going to be interminable copyright bat…

So back when Blekko was a consumer search engine we could 100% figure out who owned content on sites we crawled often. And even when we didn't we could often guess correctly more often than not based on the domain registration dates. (not to mention registry owners). That is because few people who rip off content rip off just one web site, they will rip off dozens of web sites and they will all share the same AdSense…

Or, more likely, they can't get involved for legal reasons. If they took steps to block the easy stuff, an arms race would ensue, and the content providers would never be satisfied with the performance being provided for free by Google. The content providers would always demand stricter enforcement, and could threaten to sue for copyright infringement regardless of merit.

Re: Google doesn't recognise or penalise stolen content

#63
post #62

Earlier quoted context omitted.

So back when Blekko was a consumer search engine we could 100% figure out who owned content on sites we crawled often. And even when we didn't we could often guess correctly more often than not based on the domain registration dates. (not to mention registry owners). That is because few people who rip off content rip off just one web site, they will rip off dozens of web sites and they will all share the same AdSense…

Or, more likely, they can't get involved for legal reasons. If they took steps to block the easy stuff, an arms race would ensue, and the content providers would never be satisfied with the performance being provided for free by Google. The content providers would always demand stricter enforcement, and could threaten to sue for copyright infringement regardless of merit.

I think if that were the case they would not have pushed out the "Panda" updates which penalized content farms so heavily. If their past behavior (with content farms and other "low value" content sites) is a guide they will not do anything until enough people complain about it.

In the mean time it isn't even Google's content so its not a hosting issue, they are just the "neutral" third party providing their 10 blue links (oh and supplying the advertising engine those sites are using)

Re: Google doesn't recognise or penalise stolen content

#64
post #14

This is not stealing, and even if it is illegal that is a bad way to put it. Also, google's service is primarily to the searcher so this isn't a huge issue for them.

> google's service is primarily to the searcher You are very wrong. For many years, nearly 100% of Google's revenue was from AdSense. What if someone spends days writing an article and posts it on his blog. Then, someone else copies and pastes it onto BuzzFeed, which becomes the top search result for that topic. BuzzFeed is making money that the same blogger would have made from his own content. Now, also assume Goog…

> You are very wrong. For many years, nearly 100% of Google's revenue was from AdSense.

That's incredibly untrue. A substantial portion of Google's revenue has always been and continues to be from first-party AdWords ads.

The fact that you're using BuzzFeed as an example, a firm which emphatically does not use display ads, shows how little you know about this.

Re: Google doesn't recognise or penalise stolen content

#65

I think this is pretty fair on Google's part. How could you possibly figure out who owned content? What if I published a book, it was copy-pasted in blogs, and then later I put it somewhere crawlable by Google? You certainly can't just say "first time we saw it, that's the proper owner". It would either require a massive amount of manual QA to get right (and even then, there are going to be interminable copyright bat…

one man's stolen content is another man's mirror. There are countless times where original content is region-blocked or behind a paywall, or expired, but accessible via "stolen" links.

Re: Google doesn't recognise or penalise stolen content

#66

I think this is pretty fair on Google's part. How could you possibly figure out who owned content? What if I published a book, it was copy-pasted in blogs, and then later I put it somewhere crawlable by Google? You certainly can't just say "first time we saw it, that's the proper owner". It would either require a massive amount of manual QA to get right (and even then, there are going to be interminable copyright bat…

So back when Blekko was a consumer search engine we could 100% figure out who owned content on sites we crawled often. And even when we didn't we could often guess correctly more often than not based on the domain registration dates. (not to mention registry owners). That is because few people who rip off content rip off just one web site, they will rip off dozens of web sites and they will all share the same AdSense…

Do you know if any search engine is actively filtering for this?

Re: Google doesn't recognise or penalise stolen content

#67

And they shouldn't. Thats not their job.

What is their job? I thought it was as a search engine. So, let me ask -- when people steal blog content, what are their motives for doing so? Is it to deliver that content to you, the reader? Or is it to get Google hits? How well are they preserving links, illustrations, reader comments (which are a disaster a lot of places but not all of them), an archive of other work by the same author that may be of interest? How often are they slipping undesirable things (ads that lead to sites that offer malware, for instance) alongside the content they're stealing?

If Google wants to be the best search engine possible, returning the original result for an article relevant to the user's query is a better result than returning some second-hand copy littered with low-quality ad junk. And if that's not Google job, then let me know whose job it is and I'll start using them instead.

Re: Google doesn't recognise or penalise stolen content

#68
post #24

I was huge into SEO for a few years. I try to stay out of it now, but it's worth noting that this is almost certainly due to the current algorithm's obsession with "freshness." The weaker site is ranking higher with the stolen content because their site was updated more recently. Steal some back and I bet they swap ranks again. Also, the combination of the pagerank algorithm and normal user behavior typically helps G…

the current algorithm's obsession with "freshness." That explains why I've noticed some older sites which are still around, and have plenty of detailed technical information, seem to have disappeared from the search results. Somewhat sad that the "newer is better" mentality appears to have taken over completely... if I really wanted the newest things I'd look at Google News.

I guess the problem is, there are some technical fields where old means useless. If I'm googling for Javascript libraries, hardware recommendations, or a fix to a package conflict in Ubuntu, I don't want something from 2010.

Re: Google doesn't recognise or penalise stolen content

#70

Is this any better/worse than Facebook actively trying to profit and win over users when people or organizations copy / upload / soak up views for material they did not create and don't have the rights to use? Because that's a hot-point of discussion in some creative circles as well.

Facebook's freebooting is pretty terrible. But this can destroy entire websites. The title is misleading. Google isn't just not punishing thieves, it's heavily punishing the originals. They dropped from 20th result, to 100+, because someone stole their content.

Yikes! That is much worse, at least based on your note. Do you think this is an area where the EFF could litigate on behalf of the original creators in a fraud context? Just curious, and also grateful to not be dealing with such a horrible prospect.
Post reply on HN