This is not meant to be purely controversial, but I thought long and hard about WSJ back a few months ago when HN mod (always forget his name) said to stop complaining about HN links being posted because paywalls were ok. I agree paywalls are ok. But some things are not ok. Take a look, for instance, at the WSJ.com home page with an ad blocker turned on (note all the missing letters and scrambled up titles). They wan…
If they make sending your DNA a requirement of consuming their content, then yes, you send it to them if you want their content. That's their right, as owners of something, to dictate its use. You aren't entitled to WSJ.com, NBC, or Fox.
How Google’s Web Crawler Bypasses Paywalls
171–180 of 243 posts
Re: How Google’s Web Crawler Bypasses Paywalls
#172Re: How Google’s Web Crawler Bypasses Paywalls
#173Am I alone in feeling like this is akin to a tutorial on how you can shoplift without getting caught? WSJ, for better or worse, does not want to give you content without your paying for it. If you take that content without paying, you are stealing. Just because you have figured out how to get past their security does not mean it's not stealing. (See the second precept here: https://en.wikipedia.org/wiki/Five_Precepts…
Re: How Google’s Web Crawler Bypasses Paywalls
#174Earlier quoted context omitted.
Unfortunately this will not help your defense. Andrew "Weev" Auernheimer was convicted of violating CFAA for exactly this (although the conviction was later overturned on a technicality). Again, exceeding authorized access means using your authorized access to obtain information you were not "entitled" to. So the question is not 'were you authorized' but rather it is 'were you entitled ' to that information? WTF 'ent…
> Andrew "Weev" Auernheimer was convicted of violating CFAA for exactly this (although the conviction was later overturned on a technicality). So, your example is...not an example? >WTF 'entitled' means is another question entirely No, in this case it is very clear: a request containing a particular user agent string is entitled. I have not tried this myself, but presumably you could verify that is the case by sendin…
Again I think you're confusing the fact someone could trick the server into delivering the content for free with WSJ intending to deliver their content to you for free. Since WSJ clearly intends their content to be delivered to only Googlebot for free and to users only if they pay, it is likely a jury would consider this a violation of CFAA.
A web server returning 200 OK is not ipso facto a guarantee the person making the request is not committing a crime. To give a more obvious example, if the request header contains a stolen authorization token. The law does not require the access control be non-trivial to defeat.
I don't like it, and I think the CFAA is seriously problematic, but it is the law and the Feds have been known to enforce it.
Re: How Google’s Web Crawler Bypasses Paywalls
#175If Google (or any other crawler) wanted to play nice with paywalls, they could issue a public key for their bot, and put a signature in their User Agent string that the domain could then verify. Those signatures could obviously leak, but on a per-domain basis. Perhaps the domains could have a secure way of bumping the valid key generation if they had a leak.
There are two problems with this. First, they don't want to. In fact, if a search engine can figure out that a link is going to lead to a paywall, they'll probably want to reduce the ranking of the result, because the user is not going to want results they can't actually look at. Second, it would be a massive antitrust violation because it would prevent access by competing crawlers. The only way around that is to all…
Re: How Google’s Web Crawler Bypasses Paywalls
#176Earlier quoted context omitted.
Common sense is not a set of legal procedures and rules either. The legal world cares about how the law applies to the facts of the case, not about how common sense applies. Not saying I like it.
The facts are dictated by the engineering. Is a lawyer a computer networks expert? Not by default. They will need to defer to the engineers.
However, the set of people "authorized" is not, at least not from a legal perspective. This is what the case law says. The fact that the set of people who technically _can_ access the data is different from the set of people legally authorized to access the data.
That might not be what the engineers who designed the system, run it, and produce the content intended, but that is what the law says.
It's a bummer the two disagree. But only one of the two systems put you in jail if you cross them.
You and I may wish it were otherwise, but wishing isn't going to make it so.
Re: How Google’s Web Crawler Bypasses Paywalls
#177New workaround: paste the article title into archive.is. I don't know what they're doing but they have a workaround of some sort.
Re: How Google’s Web Crawler Bypasses Paywalls
#178This is an odd debate. Let's say a restaurant declares "veterans eat free." This blog post is like a friend telling you "Hey if you tell this restaurant you're a vet they'll give you a free meal." No one said it's legal or ethical. It's lying to trick someone into giving you something at their expense. I think the relevant point, underscored by the author's last sentence, is it doesn't matter who you open a back door…
Re: How Google’s Web Crawler Bypasses Paywalls
#179Earlier quoted context omitted.
conversely, everyone should actively cloak and use random generated numbers to dynamically serve variants of their content similar to how mapmakers use trap streets. That way, like, another company wouldn't be profiting directly off of their work and threatening to sort of, destroy their entire business if you disagreed.
What? Any site can very easily not be in Google if they choose to. It's a very dumb decision for a news site, but you're free to do it.
Google has end to emd control over some users internet experience, and much of it in other cases. They own:
* 100s of thousands of servers
* domain registrar
* ~50% of web browsers in US.
* code CDN, FontService
* define web standards
* hundreds of millions of emails.
* CA implementation
* ISP infrastructure
* Develop software for a large part of thr mobile ecosystem.
* decide what you see when you go to search (most search copy google, buy results, or both)
* also many of the web beacons and advert targeting.
* oh, and the largest collection of video and images in the world.
So when you say, just do what they say or get deindexed, and you present it as if that is reasonable(not just you but the collective you) I just think I must be insane.
I mean, assuming google is good (i fo mostly) doesn't mean I would let them become the entire internet.
Real question, if google were to disappear vs. the "too big to fail banks" that would have gone under, where a case could be madr for a few certainly failing, what would have bigger impact today?
Tl;dr everyone cares about single point if failure except at the macro system level: finance, banking, healthcare, etc
Re: How Google’s Web Crawler Bypasses Paywalls
#180I'm pretty sure Google will soon stop indexing WSJ. Why index something if the vast majority of users cannot access the pages behind the links? EDIT: The "paste a headline into Google" trick still works for me, though. If this continues to be the case, they will keep indexing, of course.
>Why index something if the vast majority of users cannot access the pages behind the links? So people can find it? I'd be pissed if Google de-indexed something like IEEE because it has a paywall. Assuming the internet has to be freely available is a mistake. Especially with the continued growth of an adblocked internet. We could be facing an internet with significant paywalls in the future. I'd support a "free" sear…
WSJ is free to institute a full paywall and only serve snippets to Google. They might now like what it does to their rankings though.
What they cannot do is continue to sniff the UA before deciding to put up the paywall. (Though I'm still able to use the Google trick, so it seems the experiment might have ended.)