Live data from Hacker News

Web Scraping: Bypassing “403 Forbidden,” captchas, and more

sangaline.com

81–90 of 232 posts

Re: Web Scraping: Bypassing “403 Forbidden,” captchas, and more

#81
post #73
post #63

Earlier quoted context omitted.

How do you deal with the issue that most mobile apps have a baked in security key for their private API? Or am I being naive to think that most apps have that?

You capture the token along with the request on mitmproxy.

I guess I was being generous in my assumption that the apps will generate keys dynamically, making that not useful for a repeat attack as it were. I'm probably wrong though, most apps probably use a single baked in key.

Re: Web Scraping: Bypassing “403 Forbidden,” captchas, and more

#82

Earlier quoted context omitted.

>Doesn't publishing information on a public web server equals to granting authorization to download it? Why publish it otherwise? This is the argument that there is an implied license. The counterargument is that the user agreed to the Terms of Use which explicitly defined automated access as violative. In addition, these cases generally begin with a cease and desist demand letter, which explicitly informs the allege…

ToS is not a legally binding mutual agreement between the website and the scraper. Only when it's materialized in the form of C&D and it caused damage to the website. In the QVC vs Resultly, they caused a website crashed. In the Craigslist vs 3Taps, the damages were non-existant but they still won because of the C&D letters. So if a website sends you a letter to stop, do not continue scraping them. Until that happens…

>In the Craigslist vs 3Taps, the damages were non-existant but they still won because of the C&D letters.

(Technically they settled.)

Read up on that case and you'll see that even absent a C&D, the judge reasoned that Craiglist's IP ban against 3Taps was a separate incident of affirmatively communicating its intention that 3Taps refrain from accessing the site:

    > The calculus is different where a user is altogether banned from accessing a website.
    > The banned user has to follow only one, clear rule: do not access the website. The notice
    > issue becomes limited to how clearly the website owner communicates the banning. Here,
    > Craigslist affirmatively communicated its decision to revoke 3Taps’ access through its ceaseand-desist
    > letter and IP blocking efforts. 3Taps never suggests that those measures did not
    > put 3Taps on notice that Craigslist had banned 3Taps; indeed, 3Taps had to circumvent
    > Craigslist’s IP blocking measures to continue scraping, so it indisputably knew that Craigslist
    > did not want it accessing the website at all.
The judge continues to suggest that using proxies at all is atypical and may demonstrate an intention to violate the CFAA. Ruling at http://www.volokh.com/wp-content/uploads/2013/08/Order-Denyi... .

>So if a website sends you a letter to stop, do not continue scraping them. Until that happens, it's fair game. ToS is not legally binding. Those two cases were the result of damage caused to the website.

No, that's incorrect. You can believe this all you want, and I truly hope that it all goes well for you. It's totally possible that you will never piss off someone who has the resources to file a lawsuit over it. But you should know that you can be bound by a browsewrap agreement (and that it's quite easy to be bound by a clickwrap agreement, which, in practice, may not be much different and which scrapers may automatically follow (e.g., a link that says "click here to enter and agree")).

It's really going to come down to the judge's belief that the notification is adequately prominent and that a "reasonably prudent user" would be aware of the stipulations.

Quoth from Nguyen v. Barnes and Noble (citations removed):

    > where, as here, there is no evidence that the website
    > user had actual knowledge of the agreement, the validity of
    > the browsewrap agreement turns on whether the website puts
    > a reasonably prudent user on inquiry notice of the terms of
    > the contract. [...] Whether a user has inquiry notice of a 
    > browsewrap agreement, in turn, depends
    > on the design and content of the website and the agreement’s
    > webpage. Where the link
    > to a website’s terms of use is buried at the bottom of the page
    > or tucked away in obscure corners of the website where users
    > are unlikely to see it, courts have refused to enforce the
    > browsewrap agreement. [...] On the other hand, where the website contains an
    > explicit textual notice that continued use will act as a
    > manifestation of the user’s intent to be bound, courts have
    > been more amenable to enforcing browsewrap agreements.
    > [...] In short, the conspicuousness and
    > placement of the “Terms of Use” hyperlink, other notices
    > given to users of the terms of use, and the website’s general
    > design all contribute to whether a reasonably prudent user
    > would have inquiry notice of a browsewrap agreement.
Full decision at https://d3bsvxk93brmko.cloudfront.net/datastore/opinions/201... .

>STOP. SPREADING. FUD. cookiecaper. All of your comments are the same unsound legal advice yet Mozenda, Import.io, and a whole bunch of tools & service providers are humming along just fine.

I'm not giving any legal advice as I'm not a lawyer. For the third or fourth time here, this is all according to my layman's understanding. It's based on things I learned that time I had to close my business or face a lawsuit from a massive company over just such issues.

It's crucial for companies that provide scraping services to be aware of these issues and I know of at least one such company who is aware of them and who takes several precautions to provide some distance from potential legal liability, though they are still not 100% out of the woods. As is usually required in entrepreneurship, they're taking a calculated risk. Should someone wage a legal challenge against their activity, they have millions of dollars in the bank from investors who presumably have researched this and are willing to accept the cost of the potential legal liability.

I'm not saying that people shouldn't make businesses that depend on scraping data. I just think they should know what they're getting into before they do so.

You are correct that some businesses have been able to engage in such activities without being sued out of existence up to this point. Unfortunately, that doesn't mean that others will be as lucky.

I fully agree that anyone who is seriously interested/concerned about this should ask a lawyer. I certainly did. Their answers were not good news for me. Maybe they will be for you.

Re: Web Scraping: Bypassing “403 Forbidden,” captchas, and more

#83

Earlier quoted context omitted.

Not quite. https://www.law360.com/articles/757906/qvc-website-crash-sui... Basically, all damage claims are null because QVC & Resultly never entered into a mutual agreement. You can write whatever the fuck you want in your ToS but it's not law binding. > Judge Beetlestone also rejected QVC’s claims that Resultly violated the Computer Fraud and Abuse Act by knowingly and intentionally harming the retailer when Result…

>Basically, all damage claims are null because QVC & Resultly never entered into a mutual agreement. You can write whatever the fuck you want in your ToS but it's not law binding. I'm pretty sure that's what I said re: the ToS? That's only one element of the case (breach of contract). You are correct that in this case, browsewrap was not considered applicable. There have been a few other cases where it wasn't too, as…

Well the difference is pretty clear between your case and the rest. Developers scraping a website isn't going anywhere. A business reliant on scraped data is making money off of it. That will lend you in precarious situation more.

My criticism was that you mixed in service providers and tool providers that enable businesses to make money off scraped data-the vendor cannot be held responsible for misbehaving clients, the best it can do is cut them out when requested by external parties. Toyota doesn't appear as witness to vehicular man slaughter cases. It's a car to take you from A to B but it's not Toyota's fault if the customer runs over something between those two points and not it's intended design (QVC vs Resultly).

It also doesn't help that there are pathological web scrapers who simply does not have the money to do anything fruitful so they will bootstrap using any means necessary and plays the victim card when they are denied. This particular group is responsible for majority of the litigations. People who otherwise have no business by piggybacking off somebody else using brute force to bring heat to everyone involved.

Re: Web Scraping: Bypassing “403 Forbidden,” captchas, and more

#84
post #78

Earlier quoted context omitted.

A few things turn me off about Scrapy is that it feels over engineered for what it does. Why do I need an entire framework? I'm taking on technical debt to access data I don't have programmatic access to. CSS/Xpath are very fragile. You most likely will be changing them in the future.

> CSS/Xpath are very fragile. You most likely will be changing them in the future. Genuinely curious what the alternative is

I've been doing research on this but it's not clear whether this problem is a pain for enough number of businesses to justify further investments.

I often feel like web scraping is a commodity without understanding any of the inherent technological complexities and challenges.

Very discouraging field to be in, especially when people claim to have pain but are unwilling to pay very much for it or show appreciation for the effort that goes into it.

edit: thanks for the downvotes. perfect illustration of how innovation is punished and unrewarded in this field.

Re: Web Scraping: Bypassing “403 Forbidden,” captchas, and more

#85
post #78

Earlier quoted context omitted.

A few things turn me off about Scrapy is that it feels over engineered for what it does. Why do I need an entire framework? I'm taking on technical debt to access data I don't have programmatic access to. CSS/Xpath are very fragile. You most likely will be changing them in the future.

> CSS/Xpath are very fragile. You most likely will be changing them in the future. Genuinely curious what the alternative is

for unstructured data applying NLP

the other alternative is parsing schema.org schemas or other markup

Re: Web Scraping: Bypassing “403 Forbidden,” captchas, and more

#86

I'm curious what others use to scrape modern (javascript based) web applications. The old web (html and links) work fine with tools like Scrapy, but for modern applications which rely on javascript this does no longer work. For my last project I used a chrome plugin which controlled the browsers url locations and clicks. Results where transmitted to a backend server. New jobs (clicks, change urls) where retrieved fro…

We use HTMLUnit. Works pretty well. Not super fast, but you want to scrape individual sites at a moderate rate anyway

Re: Web Scraping: Bypassing “403 Forbidden,” captchas, and more

#87

Earlier quoted context omitted.

I'm afraid that as long as explicit agreement is not required to make a TOS binding we'll be dealing with this crap.

Yes, but see QVC v. Resultly , where the robots.txt was considered binding, not the human-readable TOS. We are getting small wins, but it's going to be slow going until we can get Congress to adjust both the CFAA and the Copyright Act, or until we can get SCOTUS to seriously alter the way these acts have been interpreted with reference to internet access.

If you don't mind de-cloaking, do you mind sending me an email?

Re: Web Scraping: Bypassing “403 Forbidden,” captchas, and more

#88

Note that in some places this constitutes breaking the law.

I think such laws are wrong. Scraping is what Google and Web Archive does and it serves good purposes. For example, one can make an application that compares prices for the same item at different internet shops and helps to find the cheapest offer. I don't understand what's wrong with downloading the information that is published on a public web server. That is what that server was made for in the first place. Of cou…

I disagree. The person who owns the server should get to decide who has access and under what circumstances. Joe Scraper, having invested nothing, has no claim or rights to it whatsoever.

Furthermore, the fact that the server gives 200 responses is not sufficient implied permission IF a no-scraping policy has been communicated in some other way such as robots.txt or (clearly communicated) TOS.

The techno-nihilist argument of "i'm just throwing bits on the wire, they're choosing to respond" is pretty unconvincing in a world where intent and context matters.

Re: Web Scraping: Bypassing “403 Forbidden,” captchas, and more

#89

Note that in some places this constitutes breaking the law.

I think such laws are wrong. Scraping is what Google and Web Archive does and it serves good purposes. For example, one can make an application that compares prices for the same item at different internet shops and helps to find the cheapest offer. I don't understand what's wrong with downloading the information that is published on a public web server. That is what that server was made for in the first place. Of cou…

That's fine as a discussion starter, but thinking a law is wrong isn't a reason to break it – if that's what you're suggesting.

That argument often turns into "But I probably won't get caught" which is a similarly weak defence.

Re: Web Scraping: Bypassing “403 Forbidden,” captchas, and more

#90
post #71

Better solution: pay target-site.com to start building an API for you. Pros: * You'll be working with them rather than against them. * Your solution will be far more robust. * It'll be way cheaper, supposing you account for the ongoing maintenance costs of your fragile scraper. * You're eliminating the possibility that you'll have to deal with legal antagonism * Good anti-scraper defenses are far more sophisticated t…

> pay target-site.com to start building an API for you. When has that ever worked?

Never.
Post reply on HN