Live data from Hacker News

Google doesn't recognise or penalise stolen content

pi-datametrics.com

21–30 of 80 posts

Re: Google doesn't recognise or penalise stolen content

#21
post #18
post #16

Earlier quoted context omitted.

> What word would you use instead? Infringing. (duh)

I can't tell if you're joking, but "infringe" means "to violate" which implies that there is a law or agreement that's being broken. That makes it sound like you agree with the idea that this is stealing.

Stealing means the victim doesn't have the stolen item anymore.

Re: Google doesn't recognise or penalise stolen content

#22
post #18
post #16

Earlier quoted context omitted.

> What word would you use instead? Infringing. (duh)

I can't tell if you're joking, but "infringe" means "to violate" which implies that there is a law or agreement that's being broken. That makes it sound like you agree with the idea that this is stealing.

I absolutely do not agree with the idea that this is stealing. How can it be stealing, when the owner still has the thing that was supposedly stolen?

Different circumstances, different terminology. The correct terminology (see US Title 17 or CDPA 1988) is "infringing". Anyone who insists on using the word "stolen" is signalling their ignorance of the first, most basic fact of copyright law.

Re: Google doesn't recognise or penalise stolen content

#23
post #17

Earlier quoted context omitted.

Plagiarism can also apply to copying content without using the exact same wording. In my mind, "stolen" means copying verbatim. Plagiarism comes from a word meaning "kidnapping" though, so the tone of both words is pretty similar.

Stolen implies that the original owner no longer has access to the data due to the actions of the perpetrator.

For physical property, yes. For intellectual property, it can still be stolen even if the original owner still has a copy.

e.g. The Soviet spies stole the plans for the hydrogen bomb.

Re: Google doesn't recognise or penalise stolen content

#24
I was huge into SEO for a few years. I try to stay out of it now, but it's worth noting that this is almost certainly due to the current algorithm's obsession with "freshness." The weaker site is ranking higher with the stolen content because their site was updated more recently. Steal some back and I bet they swap ranks again.

Also, the combination of the pagerank algorithm and normal user behavior typically helps Google to understand who was first and who deserves to rank higher. That is, most people don't plagiarize content, they quote it and then cite the source, which (thanks to pagerank) tends to rank the original better than sites which have plagiarized it.

Re: Google doesn't recognise or penalise stolen content

#25
post #17

Earlier quoted context omitted.

Plagiarism can also apply to copying content without using the exact same wording. In my mind, "stolen" means copying verbatim. Plagiarism comes from a word meaning "kidnapping" though, so the tone of both words is pretty similar.

Stolen implies that the original owner no longer has access to the data due to the actions of the perpetrator.

http://www.merriam-webster.com/dictionary/steal

See definition 1d.

Re: Google doesn't recognise or penalise stolen content

#26
post #25

Earlier quoted context omitted.

Stolen implies that the original owner no longer has access to the data due to the actions of the perpetrator.

http://www.merriam-webster.com/dictionary/steal See definition 1d.

So now Merriam-Webster has become the Twelfth Circuit? This is the lamest jurisdiction shopping I've seen in a while.

Re: Google doesn't recognise or penalise stolen content

#28
post #17

Earlier quoted context omitted.

Plagiarism can also apply to copying content without using the exact same wording. In my mind, "stolen" means copying verbatim. Plagiarism comes from a word meaning "kidnapping" though, so the tone of both words is pretty similar.

Stolen implies that the original owner no longer has access to the data due to the actions of the perpetrator.

The original owner no longer has access to the revenue stream so yes, the implication is correct.

Re: Google doesn't recognise or penalise stolen content

#30

I think this is pretty fair on Google's part. How could you possibly figure out who owned content? What if I published a book, it was copy-pasted in blogs, and then later I put it somewhere crawlable by Google? You certainly can't just say "first time we saw it, that's the proper owner". It would either require a massive amount of manual QA to get right (and even then, there are going to be interminable copyright bat…

Even if you don't care about trying to identify the original source of some piece of content, it seems like the content farm site which is plagiarizing is more likely to be a lower quality site than the original content producer.

The behavior does seem weird in any case, like there is a certain slot for a given piece of content, and Google is swapping different domains in and out to fill that slot. It seems like Google is actually trying to identify the original content, failing, and then actually inadvertently penalizing the original producer.

Post reply on HN