Live data from Hacker News

How Google’s Web Crawler Bypasses Paywalls

elaineou.com

211–220 of 243 posts

Re: How Google’s Web Crawler Bypasses Paywalls

#211
post #209

Earlier quoted context omitted.

Its a false choice without a conpelling alternative. Like saying, anyone upset with the status quo should vote. I was joking a bit, but I also wasn't. Google has end to emd control over some users internet experience, and much of it in other cases. They own: * 100s of thousands of servers * domain registrar * ~50% of web browsers in US. * code CDN, FontService * define web standards * hundreds of millions of emails.…

There are certainly more you omitted (similar to your code CDN for javascipt, fonts, etc.)... consider something like recaptcha. In my experience, all those requests to api.recaptcha.net get forwarded to "www.google.com" My experience has been that if a user for whatever reason cannot access the IP du jour for www.google.com (www.google.[cctld] will not suffice) then that user is prevented from using the myriad websi…

Right, that was my point. I actually did include code CDN and fonts in the original list, but regardless I think google is an awesome company but as a community it is just downright irresponsible to fork this level of control to an entity.

While I consider Snowden a proper hero it is almost a certainty that this could happen to a "friendly" entity like google. In that, the NSA likely has some top programmers who could get a job there and compromise something, learn enough info to find a vuln, or pass data out. This is of course making the massive assumption that they aren't already cooperating at a system level either voluntarily or involuntarily.

As you can see, as search deteriorates google is motivated to (in my opinion benevelontly) use any means neccessary to continue to fund their larger goals of a connected and automated techno-utopia. However, they will be tempted to leverage what amounts to almost literally 50% of the worlds thoughts to build systems to make short term profit while pressing forward.

Just a few that immeadiately come to mind:

* using their network to control an entire alt coin ecosystem

* using data trends to trade on global markets

* start a competing business and deindex or penalize a competitor.

* build skynet (kind of joke)

So basically, those scenarios are fairly suboptimal and I could certainly imagine that several thousand genius with knowledge of googles systems AND the worlds data could likely be profitable quickly.

Re: How Google’s Web Crawler Bypasses Paywalls

#213
post #169
post #126

Earlier quoted context omitted.

> Which, I believe, has been basis of computer "fraud" cases. This can't be true. Surely the argument for why, say, a WSJ-paywall-bypassing-tool causes damage (in the legal sense) to WSJ is that it allows people who would otherwise pay for content to get it for free, thus depriving WSJ of income. Moreover, I don't think prosecutors need to prove that you caused harm in order to charge you with computer fraud, since,…

> it allows people who would otherwise pay for content to get it for free But the sole existence of this trick and the person that would use it is exactly someone who would not pay, therefore your argument does not hold. And stealing it is not, it is more like listening to the outdoor rock concert beside the fence because you don't want to pay, inconvenient- sure, so plenty of people would still pay. Cloaking is agai…

I'm not arguing that bypassing the WSJ paywall causes damage to WSJ - I'm personally not sure that it does. What I'm arguing for is that if for some reason WSJ needed to prove in court that a paywall-bypassing tool causes them damage, they would use the it-deprived-us-of-potential-revenue argument that I outlined, rather than the it-forced-us-to-waste-electricity argument. You're right that the it-deprived-us-of-potential-revenue argument might not work in this case, in the sense that it is very possible that there does not exist a person who would have paid for a WSJ subscription but, because of the paywall-bypassing tool, did not.

Re: How Google’s Web Crawler Bypasses Paywalls

#214
post #27

And congratulations, you have likely just "exceeded authorized access" and committed a felony violation of the CFAA punishable by a fine or imprisonment for not more than 5 years under 18 U.S.C. § 1030(c)(2)(B)(i). From the ABA: "Exceeds authorized access is defined in the Computer Fraud and Abuse Act (CFAA) to mean "to access a computer with authorization and to use such access to obtain or alter information in the…

Google's _entire business_ is based not only on accessing servers without authorization, frequently in explicit violation of that site's terms of use because ToU boilerplate includes languages excluding all "crawlers, bots, or spiders", but also on flagrant violation of copyright law. They save the entire text of the web page on their servers when they crawl it (unlicensed copying), rehost it in Google Cache (unlicen…

Okay, there are other aspects to Google's business beside information retrieval and search, tho the point that the access is mostly unauthorized is valid. Although the damage done by this is to some extent offset by Google's status essentially as a public utility: it universally provides a social good, "search", and the tax we pay is advertising. Your frustration for the "little guys" is mostly en pointe, since the damage done is rendered then mostly to competitors, rather than to customers, and the argument, true or no, that Google's search is "lightyears" ahead of other possible offerings, to some extent offsets the damage to consumers due to loss of competition through the possibly anti-competitive practices you highlight. So it seems the balance of good is in the favor of consumers of "search". I think this is the main force, rather than any "structural obstacles" to competition, which is the cause of the persistence of the status quo in this market. The thing about this which no one seems to see is that, since everyone is taking information dishonestly, there's a huge opportunity to actually "bring it into the light", do it honestly, and strike some kind of deals with content creators.

Re: How Google’s Web Crawler Bypasses Paywalls

#215
post #27

And congratulations, you have likely just "exceeded authorized access" and committed a felony violation of the CFAA punishable by a fine or imprisonment for not more than 5 years under 18 U.S.C. § 1030(c)(2)(B)(i). From the ABA: "Exceeds authorized access is defined in the Computer Fraud and Abuse Act (CFAA) to mean "to access a computer with authorization and to use such access to obtain or alter information in the…

Google's _entire business_ is based not only on accessing servers without authorization, frequently in explicit violation of that site's terms of use because ToU boilerplate includes languages excluding all "crawlers, bots, or spiders", but also on flagrant violation of copyright law. They save the entire text of the web page on their servers when they crawl it (unlicensed copying), rehost it in Google Cache (unlicen…

If those sites don't want spiders, they can just specify that in robots.txt, which Google honors, right?

From the point of view of the law, that might not matter, but when there's a standardized way to make clear to bots that they're not welcome and you didn't bother to implement it, you'll look pretty silly if you complain.

Re: How Google’s Web Crawler Bypasses Paywalls

#216
post #86

Earlier quoted context omitted.

try deleting cookies, then hit refresh.

Not working here either. I tried wiping out cookies and cache, and Chrome is sending the right user agent (Googlebot).

Try disabling other extensions that might alter the user agent.

Re: How Google’s Web Crawler Bypasses Paywalls

#217

"Remember: Any time you introduce an access point for a trusted third party, you inevitably end up allowing access to anybody." See also: http://www.apple.com/customer-letter/ :)

Google also specifies the ip ranges of their boys; just UA checking is sloppy

Re: How Google’s Web Crawler Bypasses Paywalls

#218
post #202

Earlier quoted context omitted.

In this example, the FBI is the "trusted third party", but by giving them access, we inevitably open access for everyone, as the system is no longer strongly secure. The trusted third party in the quote isn't asking for access for everybody either, but in the end that's what happens.

Apple isn't giving access. Apple would be required (by court, unless they manage to fight this off) to install a signed custom build of the OS in order to give access to that particular device. FBI would not have this build, nor a key to create their own signed custom build.

Yup, that's the FBI's pitch. The issue is that it sets a legal precedent as well as potentially leaking a backdoored iOS to the world. Yes I know "signed for a specific device", best of luck with that.

Re: How Google’s Web Crawler Bypasses Paywalls

#219
post #82

If Google (or any other crawler) wanted to play nice with paywalls, they could issue a public key for their bot, and put a signature in their User Agent string that the domain could then verify. Those signatures could obviously leak, but on a per-domain basis. Perhaps the domains could have a secure way of bumping the valid key generation if they had a leak.

Google (and every other major search engine) already provide a way, i.e. reverse DNS lookup, to authentic bot ownership:

https://support.google.com/webmasters/answer/80553?hl=en

AFAIK no content provider actually does this check though.

Re: How Google’s Web Crawler Bypasses Paywalls

#220
post #153

Earlier quoted context omitted.

I believe this refers to the Weev case.

It may, but weev is probably not the only person whose been put away for that. This type of activity is the basis for Google and many other tech startups. I hope they catch Larry and Sergei soon, they've been on the lam for almost two decades!

But that one is for "greater public good" you see .... ;)
Post reply on HN