Live data from Hacker News

How Google’s Web Crawler Bypasses Paywalls

elaineou.com

201–210 of 243 posts

Re: How Google’s Web Crawler Bypasses Paywalls

#201

"Remember: Any time you introduce an access point for a trusted third party, you inevitably end up allowing access to anybody." See also: http://www.apple.com/customer-letter/ :)

I was going to post here that there were ways around that; proper security, cryptographic access control... And then I saw the light. ;)

Re: How Google’s Web Crawler Bypasses Paywalls

#202
post #186

Earlier quoted context omitted.

Except of course what FBI is proposing wouldn't give access to "anybody". Just unfettered access for FBI, CIA and NSA by way of gag orders and national security letters, forcing Apple to break security of their hardware and not speak about it publicly. Of note is the fact that it wouldn't even give direct access to FBI, since they don't have the firmware keys (yet).

In this example, the FBI is the "trusted third party", but by giving them access, we inevitably open access for everyone, as the system is no longer strongly secure. The trusted third party in the quote isn't asking for access for everybody either, but in the end that's what happens.

Apple isn't giving access. Apple would be required (by court, unless they manage to fight this off) to install a signed custom build of the OS in order to give access to that particular device. FBI would not have this build, nor a key to create their own signed custom build.

Re: How Google’s Web Crawler Bypasses Paywalls

#203
post #197

Earlier quoted context omitted.

> Andrew "Weev" Auernheimer was convicted of violating CFAA for exactly this (although the conviction was later overturned on a technicality). So, your example is...not an example? >WTF 'entitled' means is another question entirely No, in this case it is very clear: a request containing a particular user agent string is entitled. I have not tried this myself, but presumably you could verify that is the case by sendin…

> So, your example is...not an example? At least in the US, the law doesn't work that way. Decisions will quite often cite some other similar case which reached the opposite conclusion , but under different circumstances, because that other case's decision says something like "X, if it weren't for Y" or "Fortunately for the defendant, they didn't Z, so not X", or something. That isn't binding precedent for the judge…

We aren't talking about user names and passwords, but user agent strings.

The other decision was vacated. The jury was not appropriate and their decision is irrelevant.

Re: How Google’s Web Crawler Bypasses Paywalls

#204

If this is true, what WSJ is doing is called "cloaking" and should cause it to get de-indexed: https://support.google.com/webmasters/answer/66355?hl=en

conversely, everyone should actively cloak and use random generated numbers to dynamically serve variants of their content similar to how mapmakers use trap streets. That way, like, another company wouldn't be profiting directly off of their work and threatening to sort of, destroy their entire business if you disagreed.

Ha yea, why let Google have the power to destroy your business when you can burn it to the ground yourself!

Re: How Google’s Web Crawler Bypasses Paywalls

#205
post #174

Earlier quoted context omitted.

> Andrew "Weev" Auernheimer was convicted of violating CFAA for exactly this (although the conviction was later overturned on a technicality). So, your example is...not an example? >WTF 'entitled' means is another question entirely No, in this case it is very clear: a request containing a particular user agent string is entitled. I have not tried this myself, but presumably you could verify that is the case by sendin…

It's the best example we got. The case was overturned (after he spent quite some time in federal prison) not because it was found that he didn't violate the CFAA but because the charges were brought in the wrong jurisdiction. Again I think you're confusing the fact someone could trick the server into delivering the content for free with WSJ intending to deliver their content to you for free. Since WSJ clearly intends…

As far as the law is concerned, he did not violate anything. You do not have to prove yourself innocent, the burden is on the prosecutor to prove a violation. In the example you cite, no violation has been shown.

It is not at all clear that WSJ intends Googlebot to get their content for free while others must pay. Thus is actually against Google's policies, which would actually call into question whether WSJ's behavior is felonius. WSJ may not be entitled to be incorporated into Google's index, yet they are manipulating the Googlebot to the contrary.

Re: How Google’s Web Crawler Bypasses Paywalls

#206
post #65

Earlier quoted context omitted.

> If they allowed all bots but disallowed all non-bots, that would raise questions of what defines a "bot". Maybe, but I think its a pretty easy distinction. They aren't even allowing all bots - they're allowing a white list of them. You're not just writing your own bot to get around it, you're pretending to be someone else's bot. > But in practice, it seems they favor certain bots. Is it ok that the WSJ lets Google…

> If you view their website as a physical store / private establishment, then I assume that they have every right to establish who has access to what and under what conditions. Not true. There are very specific laws about not being able to discriminate against protected classes of people. Bars have to serve minorities, bakeries have to cater to same sex marriages, etc. Where you draw the line of legislated equality w…

That's fair, though private establishments are still allowed to distinguish based on other criteria such as membership, partnership, etc. I didn't meant to be super specific. I'm merely pointing out that they are allowed to establish barriers to entry.

Re: How Google’s Web Crawler Bypasses Paywalls

#207

I was under the impression that the "hack" whereby you searched for the article on Google and clicked through to that article (effectively skipping over the paywall) was a demand of Google's and not an oversight by the paywalled website. I thought that google deemed providing search results which were behind paywalls as a "bad experience" for their search users, and would penalize websites for doing so. Is this no lo…

Google doesn't demand anything. If your paywalled website is not accessible by Google's crawler, then Google will not index it. Publishers want Google to index their pages and drive potential paying visitors, which is why they open the loophole themselves. For the second point, Google does require that publishers specify "registration required" in their sitemap.

If you're showing Googlebot one thing, and visitors who visit your website through google another, that's essentially "cloaking" (a blackhat SEO technique). At least it used to be.

Re: How Google’s Web Crawler Bypasses Paywalls

#209
post #47

Earlier quoted context omitted.

What? Any site can very easily not be in Google if they choose to. It's a very dumb decision for a news site, but you're free to do it.

Its a false choice without a conpelling alternative. Like saying, anyone upset with the status quo should vote. I was joking a bit, but I also wasn't. Google has end to emd control over some users internet experience, and much of it in other cases. They own: * 100s of thousands of servers * domain registrar * ~50% of web browsers in US. * code CDN, FontService * define web standards * hundreds of millions of emails.…

There are certainly more you omitted (similar to your code CDN for javascipt, fonts, etc.)... consider something like recaptcha.

In my experience, all those requests to api.recaptcha.net get forwarded to "www.google.com"

My experience has been that if a user for whatever reason cannot access the IP du jour for www.google.com (www.google.[cctld] will not suffice) then that user is prevented from using the myriad websites that rely on recpatcha.net.

Now, I could be wrong and maybe there is something I am missing, but in my experience this is a sad state of centralization and reliance by websites on Google. Quite brittle.

Re: How Google’s Web Crawler Bypasses Paywalls

#210
post #61

Earlier quoted context omitted.

So? Is this a "two wrongs make a right?" More importantly, is violating "Google's rules" suddenly a violation of law?

I don't think its wrong to browse a site's content that they have advertised as publically available? If they are advertising incorrectly, they should fix that.

False advertising is also illegal
Post reply on HN