Live data from Hacker News

Google doesn't recognise or penalise stolen content

pi-datametrics.com

41–50 of 80 posts

Re: Google doesn't recognise or penalise stolen content

#41
post #28

Earlier quoted context omitted.

Stolen implies that the original owner no longer has access to the data due to the actions of the perpetrator.

The original owner no longer has access to the revenue stream so yes, the implication is correct.

Then, say "potential revenue" was stolen, not content.

Re: Google doesn't recognise or penalise stolen content

#42

I think this is pretty fair on Google's part. How could you possibly figure out who owned content? What if I published a book, it was copy-pasted in blogs, and then later I put it somewhere crawlable by Google? You certainly can't just say "first time we saw it, that's the proper owner". It would either require a massive amount of manual QA to get right (and even then, there are going to be interminable copyright bat…

While most of what you say is true, they could at the very least reject AdSense applicants based on how often they copy-paste content from established publishers. I’m sure they have the means to figure something like that out.

Re: Google doesn't recognise or penalise stolen content

#44
post #25

Earlier quoted context omitted.

http://www.merriam-webster.com/dictionary/steal See definition 1d.

A different angle is to look at what is accepted in a court of law. Laws do not consider copying/infringement to be stealing. I know some people who say gunna (I was gunna do it). However, that doesn't make it correct. So, usage is not a measure of validity either. Most online dictionaries exclude steal/theft being linked to infringement/copyright.

EDIT: Look at OED's additions list from June 2015, specifically under 'g': http://public.oed.com/the-oed-today/recent-updates-to-the-oe...

'gunna'. Is it still wrong?

--original comment--

Common usage does eventually lead to validity, though. Language is not static and evolves through usage. Besides, how are you going to measure validity? Is the OED the sole arbiter of what's "correct"? Common usage and mutual understanding is a great way to determine what's "correct" in a language.

I won't debate the legal usage of the word, as there was no mention of legal interpretation by the courts in the previous comments. However you are correct from a legal perspective where words have specific meanings.

Re: Google doesn't recognise or penalise stolen content

#45
post #23

Earlier quoted context omitted.

Stolen implies that the original owner no longer has access to the data due to the actions of the perpetrator.

For physical property, yes. For intellectual property, it can still be stolen even if the original owner still has a copy. e.g. The Soviet spies stole the plans for the hydrogen bomb.

Ah, so you mean the actual blueprints? And the poor Americans had no copies? How stupid of them!

Re: Google doesn't recognise or penalise stolen content

#46
post #23

Earlier quoted context omitted.

Stolen implies that the original owner no longer has access to the data due to the actions of the perpetrator.

For physical property, yes. For intellectual property, it can still be stolen even if the original owner still has a copy. e.g. The Soviet spies stole the plans for the hydrogen bomb.

So... copied.

Re: Google doesn't recognise or penalise stolen content

#47
post #11
post #6

There is no such thing as "Stolen" content.

If that's true, there's also no such thing as "stealing" at all. Consider a novelist who works for 10 years on her novel. A hacker steals the document from her computer and publishes it online under his own name. He makes $100M. Is it wrong for the novelist to feel like someone stole from her? What word would you use instead?

How are people still arguing about this?

Saying that someone is "stealing" when they infringe copyright is like saying someone is "killing you" when they present convincing arguments against your cause. It isn't literally stealing or killing, it's an exaggeration made for emphasis.

The reason there is so much contention is that a) the same language has been extremely common among hysterical content industry lobbyists who insist that it is literally stealing, and b) stealing and copyright infringement are both unlawful (and therefore more easily confused) even though there remains a meaningful distinction between stealing and copying.

But that distinction is very important in practice because we can't treat stealing and infringement the same. If you don't like someone's speech you can't be allowed to steal any of their webservers but you have to be allowed to copy some of their work in order to effectively criticize them.

Re: Google doesn't recognise or penalise stolen content

#48

I think this is pretty fair on Google's part. How could you possibly figure out who owned content? What if I published a book, it was copy-pasted in blogs, and then later I put it somewhere crawlable by Google? You certainly can't just say "first time we saw it, that's the proper owner". It would either require a massive amount of manual QA to get right (and even then, there are going to be interminable copyright bat…

Also all of this is assuming that any duplicated content is inherently stolen, when it could in fact be public domain, fair use, legitimately licensed, distributed under Creative Commons, etc.

True, but in these cases I would still want original/canonical/fastest/best source first while the others are probably only valuable as backups.

Re: Google doesn't recognise or penalise stolen content

#49
post #24

I was huge into SEO for a few years. I try to stay out of it now, but it's worth noting that this is almost certainly due to the current algorithm's obsession with "freshness." The weaker site is ranking higher with the stolen content because their site was updated more recently. Steal some back and I bet they swap ranks again. Also, the combination of the pagerank algorithm and normal user behavior typically helps G…

> the current algorithm's obsession with "freshness."

Which is how Google makes blogspam such a good business to be in, even if your content is inferior to the post you used for "research".

Re: Google doesn't recognise or penalise stolen content

#50
post #30

I think this is pretty fair on Google's part. How could you possibly figure out who owned content? What if I published a book, it was copy-pasted in blogs, and then later I put it somewhere crawlable by Google? You certainly can't just say "first time we saw it, that's the proper owner". It would either require a massive amount of manual QA to get right (and even then, there are going to be interminable copyright bat…

Even if you don't care about trying to identify the original source of some piece of content, it seems like the content farm site which is plagiarizing is more likely to be a lower quality site than the original content producer. The behavior does seem weird in any case, like there is a certain slot for a given piece of content, and Google is swapping different domains in and out to fill that slot. It seems like Goog…

Well, Google has already indexed a new article x. When article y appears, and Google sees that y is an almost verbatim repeat of x, it shouldn't be that hard to figure out that article x is the original, should it? Especially if they both have time/date stamps....
Post reply on HN