Live data from Hacker News

How Google’s Web Crawler Bypasses Paywalls

elaineou.com

81–90 of 243 posts

Re: How Google’s Web Crawler Bypasses Paywalls

#81
post #61
post #50

Earlier quoted context omitted.

WSJ is violating Google's rules referenced here: https://support.google.com/news/publisher/answer/40543?hl=en

So? Is this a "two wrongs make a right?" More importantly, is violating "Google's rules" suddenly a violation of law?

I don't think its wrong to browse a site's content that they have advertised as publically available?

If they are advertising incorrectly, they should fix that.

Re: How Google’s Web Crawler Bypasses Paywalls

#82
If Google (or any other crawler) wanted to play nice with paywalls, they could issue a public key for their bot, and put a signature in their User Agent string that the domain could then verify.

Those signatures could obviously leak, but on a per-domain basis. Perhaps the domains could have a secure way of bumping the valid key generation if they had a leak.

Re: How Google’s Web Crawler Bypasses Paywalls

#83
post #76

Earlier quoted context omitted.

Interestingly, the CFAA does not define the term "without authorization" however it does define "exceeding authorization" exactly as I quoted above; - to access a computer with authorization and to use such access to obtain or alter information in the computer that the accesser is not entitled so to obtain or alter. So arguing User-agent is not an authorization mechanism probably won't help you, because exceeding aut…

If the computer on the public Internet responds to a standard HTTP request, it has explicitly authorized my access to whatever information it sent me.

Unfortunately this will not help your defense. Andrew "Weev" Auernheimer was convicted of violating CFAA for exactly this (although the conviction was later overturned on a technicality).

Again, exceeding authorized access means using your authorized access to obtain information you were not "entitled" to. So the question is not 'were you authorized' but rather it is 'were you entitled' to that information? WTF 'entitled' means is another question entirely, but likely it is in the eye of the beholder. A jury decided Weev was not 'entitled' to the email addresses he downloaded from AT&T, and it's safe to assume we are not 'entitled' to free access to WSJ's content. So I would not rest your hopes on the "200 OK".

Re: How Google’s Web Crawler Bypasses Paywalls

#85
post #76

Earlier quoted context omitted.

Interestingly, the CFAA does not define the term "without authorization" however it does define "exceeding authorization" exactly as I quoted above; - to access a computer with authorization and to use such access to obtain or alter information in the computer that the accesser is not entitled so to obtain or alter. So arguing User-agent is not an authorization mechanism probably won't help you, because exceeding aut…

If the computer on the public Internet responds to a standard HTTP request, it has explicitly authorized my access to whatever information it sent me.

Your request wouldn't be 'standard', it would be deliberately malformed to bypass a paywall. You guys are performing linguistic gymnastics to get around that fact.

Re: How Google’s Web Crawler Bypasses Paywalls

#86

My Windows anti-virus deletes the linked sample code automatically upon download, marking it as "Trojan:Win32/Spursint.A". Did anyone have the same experience? (I was actually more interested in using it as a template for writing a simple Chrome extension.)

Yep. I then pasted it but it didn't work on wsj.com. Oh well.

try deleting cookies, then hit refresh.

Re: How Google’s Web Crawler Bypasses Paywalls

#87
post #58

Earlier quoted context omitted.

Indexing is fine, a great feature would be if Google was able to show it only to the user that can access it.

Why should Google manage WSJ's paywall?

It doesn't. It 1) penalize WSJ, 2) personalize to WSJ subscribed Google users.

Re: How Google’s Web Crawler Bypasses Paywalls

#88
post #51

Earlier quoted context omitted.

It's intent that matters. Setting user-agent in order to properly render a page is legal. Setting a user-agent string to gain access to otherwise unauthorized content is probably not.

User-agent string is not an authorization mechanism.

Good luck arguing that in court. Esp against the multiple lawyers a corporation will be able to afford.

If you don't get it, court doesn't care about what's reasonableness, technically correctness, etc. Only if your lawyers can convince jury/judge. Twinky made me crazy, It the gloves don't fit... and so forth.

Re: How Google’s Web Crawler Bypasses Paywalls

#89

Earlier quoted context omitted.

I'm surprised that most online papers won't sell you one day's worth online for a buck or so. Like buying a real newspaper. They all seem to want to sell subscriptions, which are perpetual and probably difficult to cancel..

Even a dollar a day. I can pay Netflix $10 a month and stream unlimited HD video, but wsj wants $30 to read the first few paragraphs of a few articles a day? The pricing here is much too aggressive

Videos are more relevant to rewatch than articles are to reread. One needs to output more articles than videos, because articles must be "fresh" or you'll lose an audience. Nobody is printing 40 year old news - people are still watching 40 year old movies.

Playing devil's advocate here. Pricing for many online goods is almost completely arbitrary and varies with little accord to service/product quality or even what that service provides.

Another related example of arbitrary pricing: people will pay $2 for a soda from a vending machine but won't pay $1 for a useful app on their phone.

There's something going on there... The sooner the $1 app's figure out what makes people buy $2 sodas is the day they become rich. And the sooner content providers figure out why people will pay $20/month to stream media (Let's say $10 to Netflix, $10 to Spotify or something) and charge people $20/mo for their articles... things will turn around for them.

Re: How Google’s Web Crawler Bypasses Paywalls

#90
post #20

Earlier quoted context omitted.

A App Engine wouldn't have a IP with a reverse DNS *.googlebot.com, would it?

Nope, but it does resolve to something .google and if you check netblocks.google.com it will appear there so they might not be limiting it to googlebot only at this point. https://support.google.com/a/answer/60764?hl=en

_netblocks.google.com includes only the servers that handle Gmail and corporate email.
Post reply on HN