Live data from Hacker News

How Google’s Web Crawler Bypasses Paywalls

elaineou.com

61–70 of 243 posts

Re: How Google’s Web Crawler Bypasses Paywalls

#61
post #50
post #24

Am I alone in feeling like this is akin to a tutorial on how you can shoplift without getting caught? WSJ, for better or worse, does not want to give you content without your paying for it. If you take that content without paying, you are stealing. Just because you have figured out how to get past their security does not mean it's not stealing. (See the second precept here: https://en.wikipedia.org/wiki/Five_Precepts…

WSJ is violating Google's rules referenced here: https://support.google.com/news/publisher/answer/40543?hl=en

So? Is this a "two wrongs make a right?"

More importantly, is violating "Google's rules" suddenly a violation of law?

Re: How Google’s Web Crawler Bypasses Paywalls

#62
post #31

Earlier quoted context omitted.

they're probably doing some sort of a/b testing by selectively letting some clicks through

This is basically true ^^

Any idea what the deal is with SEO impacts of WSJ taking the idea of blocking everyone who isn't a google bot?

Re: How Google’s Web Crawler Bypasses Paywalls

#63
post #51

Earlier quoted context omitted.

User agent strings have a long history of being intentionally misleading. IE 11 claims to be "Mozilla/5.0". Chrome claims to be "Safari/537.36". The User-Agent string is all lies, and has been ever since the first site started doing UA sniffing.

It's intent that matters. Setting user-agent in order to properly render a page is legal. Setting a user-agent string to gain access to otherwise unauthorized content is probably not.

User-agent string is not an authorization mechanism.

Re: How Google’s Web Crawler Bypasses Paywalls

#64
post #32

Earlier quoted context omitted.

Who has? Surely not the person who wrote the tutorial.

Under CFAA, I don't know. The DMCA may have some problems with that blog post, though. And by "may", I do mean "may". I don't know. But it's at least possible.

That's a great point, it's likely an illegal DMCA circumvention device too!

  No person shall manufacture, import, offer to the public, provide, or otherwise traffic
  in any technology, product, service, device, component, or part thereof, that—

  (A) is primarily designed or produced for the purpose of circumventing a technological
      measure that effectively controls access to a work protected under this title;

  (B) has only limited commercially significant purpose or use other than to circumvent a 
      technological measure that effectively controls access to a work protected under
      this title; or

  (C) is marketed by that person or another acting in concert with that person with that
      person’s knowledge for use in circumventing a technological measure that effectively
      controls access to a work protected under this title.
I mean, obviously it's all quite ridiculous, but also deadly serious at the same time :-(

Re: How Google’s Web Crawler Bypasses Paywalls

#65
post #37
post #24

Am I alone in feeling like this is akin to a tutorial on how you can shoplift without getting caught? WSJ, for better or worse, does not want to give you content without your paying for it. If you take that content without paying, you are stealing. Just because you have figured out how to get past their security does not mean it's not stealing. (See the second precept here: https://en.wikipedia.org/wiki/Five_Precepts…

Legally it might, I don't know. Morally, I'm not sure. 1. If they allowed all bots but disallowed all non-bots, that would raise questions of what defines a "bot". Couldn't I write a personal bot that fetches the story for me? As a browser addon, even? 2. It's even more complex since allowing bots means they allow tools that provide the information to third parties, as the bots are not intended for private use by the…

> If they allowed all bots but disallowed all non-bots, that would raise questions of what defines a "bot".

Maybe, but I think its a pretty easy distinction. They aren't even allowing all bots - they're allowing a white list of them. You're not just writing your own bot to get around it, you're pretending to be someone else's bot.

> But in practice, it seems they favor certain bots. Is it ok that the WSJ lets Google do things Google's competitors cannot?

That's the really important question. I personally have no context for answering except to say that I can see both sides argued. If you view their website as a physical store / private establishment, then I assume that they have every right to establish who has access to what and under what conditions.

Of course, that hampers a lot of legitimate use cases along the way.

Re: How Google’s Web Crawler Bypasses Paywalls

#66
post #27

And congratulations, you have likely just "exceeded authorized access" and committed a felony violation of the CFAA punishable by a fine or imprisonment for not more than 5 years under 18 U.S.C. § 1030(c)(2)(B)(i). From the ABA: "Exceeds authorized access is defined in the Computer Fraud and Abuse Act (CFAA) to mean "to access a computer with authorization and to use such access to obtain or alter information in the…

User agent strings have a long history of being intentionally misleading. IE 11 claims to be "Mozilla/5.0". Chrome claims to be "Safari/537.36". The User-Agent string is all lies, and has been ever since the first site started doing UA sniffing.

A decent Friday afternoon read on that topic:

http://webaim.org/blog/user-agent-string-history/

Re: How Google’s Web Crawler Bypasses Paywalls

#70
post #58

I'm pretty sure Google will soon stop indexing WSJ. Why index something if the vast majority of users cannot access the pages behind the links? EDIT: The "paste a headline into Google" trick still works for me, though. If this continues to be the case, they will keep indexing, of course.

Indexing is fine, a great feature would be if Google was able to show it only to the user that can access it.

Why should Google manage WSJ's paywall?
Post reply on HN