Live data from Hacker News

How Google’s Web Crawler Bypasses Paywalls

elaineou.com

101–110 of 243 posts

Re: How Google’s Web Crawler Bypasses Paywalls

#101
I thought Google specifically disallowed returning different pages based on User-Agent targetting googlebot, and this included paywalls.

Are they running afoul of Google policies and going to get pinged by Google?

I can't find the text from Google now (when can you ever find any docs at google?), but I am very certain I remember reading from them that you may not return different content to GoogleBot based on User-Agent.

Re: How Google’s Web Crawler Bypasses Paywalls

#103
post #78
post #51

Earlier quoted context omitted.

It's intent that matters. Setting user-agent in order to properly render a page is legal. Setting a user-agent string to gain access to otherwise unauthorized content is probably not.

In that case, I'm just going to set my User-Agent permanently to Google's crawler for fun and see how the web renders, in general. Just for fun, because I want to see what the world looks like from a different perspective. That's my intent. I've stated it. Now that my user agent has been changed, I'm going to go grab lunch, carry on with life, maybe check out some cat pictures and then maybe some news sites over tea…

In my opinion the WSJ has a major rendering bug that affects everyone except for the google search crawler, and I selectively set my User Agent to get around said bug.

Re: How Google’s Web Crawler Bypasses Paywalls

#104
post #37
post #24

Am I alone in feeling like this is akin to a tutorial on how you can shoplift without getting caught? WSJ, for better or worse, does not want to give you content without your paying for it. If you take that content without paying, you are stealing. Just because you have figured out how to get past their security does not mean it's not stealing. (See the second precept here: https://en.wikipedia.org/wiki/Five_Precepts…

Legally it might, I don't know. Morally, I'm not sure. 1. If they allowed all bots but disallowed all non-bots, that would raise questions of what defines a "bot". Couldn't I write a personal bot that fetches the story for me? As a browser addon, even? 2. It's even more complex since allowing bots means they allow tools that provide the information to third parties, as the bots are not intended for private use by the…

[deleted]

Re: How Google’s Web Crawler Bypasses Paywalls

#106

Earlier quoted context omitted.

As you've indicated, a request can only be considered malformed if it doesn't conform to the standards. Here is what the relevant RFC[0] has to say about the User-Agent header: > Likewise, implementations are encouraged not to use the product tokens of other implementations in order to declare compatibility with them, as this circumvents the purpose of the field. If a user agent masquerades as a different user agent,…

Unless you were being tongue-in-cheek, you've completely sidestepped the intent of my comment and continued with the linguistic silliness. Obtaining paywall-protected content by faking your user agent to purport yourself to be a Google Crawler is quite clearly fraudulent. This isn't a point for debate. PS. To play along with the linguistic theme, can you provide a source for the definition of a malformed request? My…

I wasn't playing a linguistic game. You mentioned standards and malformed requests, and I pointed out that you misused those terms. I am not bothering to track down definitions for you to play, as you say, linguistic games. You are free to find something that proves my position wrong, as I used the RFC to cite my position as correct.

If you meant to put forth that faking a user agent is a technique to exceed authorization, that's fine, and I'm glad to have helped you clarify it. Just be clear, it's not what you said with your detour into malformed requests.

Re: How Google’s Web Crawler Bypasses Paywalls

#107
So does HN now choose to not post articles from the WSJ? I was comfortable with the "google it" trick, and frankly was a little annoyed with constant "paywall, wah!" comments when what should be by now a well-known workaround was available. But that workaround no longer works.

Re: How Google’s Web Crawler Bypasses Paywalls

#108
post #27

And congratulations, you have likely just "exceeded authorized access" and committed a felony violation of the CFAA punishable by a fine or imprisonment for not more than 5 years under 18 U.S.C. § 1030(c)(2)(B)(i). From the ABA: "Exceeds authorized access is defined in the Computer Fraud and Abuse Act (CFAA) to mean "to access a computer with authorization and to use such access to obtain or alter information in the…

User agent strings have a long history of being intentionally misleading. IE 11 claims to be "Mozilla/5.0". Chrome claims to be "Safari/537.36". The User-Agent string is all lies, and has been ever since the first site started doing UA sniffing.

Pretty obvious weak move restricting by UA$.

Allow by IP range? You can probably find a somewhat accurate range for Google and whoever's crawlers.

Re: How Google’s Web Crawler Bypasses Paywalls

#109
post #51

Earlier quoted context omitted.

User agent strings have a long history of being intentionally misleading. IE 11 claims to be "Mozilla/5.0". Chrome claims to be "Safari/537.36". The User-Agent string is all lies, and has been ever since the first site started doing UA sniffing.

It's intent that matters. Setting user-agent in order to properly render a page is legal. Setting a user-agent string to gain access to otherwise unauthorized content is probably not.

Unauthorized? If you put up a sign and tell people they have to pay to look at it, is it illegal to look at it and not pay? This should be a rhetorical question.

Re: How Google’s Web Crawler Bypasses Paywalls

#110
post #27

And congratulations, you have likely just "exceeded authorized access" and committed a felony violation of the CFAA punishable by a fine or imprisonment for not more than 5 years under 18 U.S.C. § 1030(c)(2)(B)(i). From the ABA: "Exceeds authorized access is defined in the Computer Fraud and Abuse Act (CFAA) to mean "to access a computer with authorization and to use such access to obtain or alter information in the…

How does that law apply to foreigners?
Post reply on HN