Live data from Hacker News

How Google’s Web Crawler Bypasses Paywalls

elaineou.com

21–30 of 243 posts

Re: How Google’s Web Crawler Bypasses Paywalls

#21
post #3

Earlier quoted context omitted.

Did you read the article. It talks about how that trick no longer works on a lot of sites because they are now checking User-Agent strings too.

Actually I noticed sites have simply changed policies -- if you're a regular visitor your cookies will identify you and block content. The Incognito mode trick works for WSJ and others that would still check the referrer header. Allowing Googlebot access and checking the referrer header are two different things. Also, Google has published IP addresses it uses, so this extension might not last long...

> Also, Google has published IP addresses it uses, so this extension might not last long...

They do not [1], but you can find out by doing a reverse DNS query.

> "Google doesn't post a public list of IP addresses for webmasters to whitelist"

[1] https://support.google.com/webmasters/answer/80553?hl=en

Re: How Google’s Web Crawler Bypasses Paywalls

#22
post #20

Earlier quoted context omitted.

Seems to work if you deploy a proxy on Google's app engine and use it to access WSJ ;)

A App Engine wouldn't have a IP with a reverse DNS *.googlebot.com, would it?

Nope, but it does resolve to something .google and if you check netblocks.google.com it will appear there so they might not be limiting it to googlebot only at this point.

https://support.google.com/a/answer/60764?hl=en

Re: How Google’s Web Crawler Bypasses Paywalls

#23
post #17

I just tried clicking on "Harper Lee, Author of ‘To Kill a Mockingbird,’ Dies at Age 89" from wsj.com's homepage and got the paywall. I then pasted the headline into google and clicked on it from Google results and did not get hit by the paywall.

I also checked the paywall and the old trick still works for me. Odd.

Re: How Google’s Web Crawler Bypasses Paywalls

#24
Am I alone in feeling like this is akin to a tutorial on how you can shoplift without getting caught? WSJ, for better or worse, does not want to give you content without your paying for it. If you take that content without paying, you are stealing. Just because you have figured out how to get past their security does not mean it's not stealing.

(See the second precept here: https://en.wikipedia.org/wiki/Five_Precepts)

Re: How Google’s Web Crawler Bypasses Paywalls

#25

If this is true, what WSJ is doing is called "cloaking" and should cause it to get de-indexed: https://support.google.com/webmasters/answer/66355?hl=en

conversely, everyone should actively cloak and use random generated numbers to dynamically serve variants of their content similar to how mapmakers use trap streets. That way, like, another company wouldn't be profiting directly off of their work and threatening to sort of, destroy their entire business if you disagreed.

Re: How Google’s Web Crawler Bypasses Paywalls

#27
And congratulations, you have likely just "exceeded authorized access" and committed a felony violation of the CFAA punishable by a fine or imprisonment for not more than 5 years under 18 U.S.C. § 1030(c)(2)(B)(i).

From the ABA: "Exceeds authorized access is defined in the Computer Fraud and Abuse Act (CFAA) to mean "to access a computer with authorization and to use such access to obtain or alter information in the computer that the accesser is not entitled so to obtain or alter."

To prove you have committed this terrible felony, the FBI will now demand that Apple assist in disabling the secure enclave of your device in order to access your browser history. But remember, they only need to do this because they aren't allow to MITM all TLS and "acquire" -- not "collect" -- every HTTP request your machine ever makes.

Re: How Google’s Web Crawler Bypasses Paywalls

#28
Correct me if I'm wrong, but wasn't there a long standing Google's policy that the version of the page served to their crawler must also be publicly accessible. That would then be the reason why WSJ articles were accessible through the paste-into-google trick, rather than because WSJ was incompetent and failed to "fix" the bypass.

So does it mean that Google will no longer index full WSJ articles or does it mean a change in the Google's policy?

Re: How Google’s Web Crawler Bypasses Paywalls

#29
post #27

And congratulations, you have likely just "exceeded authorized access" and committed a felony violation of the CFAA punishable by a fine or imprisonment for not more than 5 years under 18 U.S.C. § 1030(c)(2)(B)(i). From the ABA: "Exceeds authorized access is defined in the Computer Fraud and Abuse Act (CFAA) to mean "to access a computer with authorization and to use such access to obtain or alter information in the…

Who has? Surely not the person who wrote the tutorial.

Re: How Google’s Web Crawler Bypasses Paywalls

#30
post #24

Am I alone in feeling like this is akin to a tutorial on how you can shoplift without getting caught? WSJ, for better or worse, does not want to give you content without your paying for it. If you take that content without paying, you are stealing. Just because you have figured out how to get past their security does not mean it's not stealing. (See the second precept here: https://en.wikipedia.org/wiki/Five_Precepts…

[deleted]
Post reply on HN