Live data from Hacker News

9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

cdn.ca9.uscourts.gov

151–160 of 293 posts

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#151
post #131

Earlier quoted context omitted.

When it comes to physical properties there's a huge difference between reading a banner posted in a street and entering the property to read some secret data: you have to be in different locations. That's why your analogy is completely faulty. When it comes to PUBLIC data in a website there's no difference. How would I know I'm authorized, implicitly or explicitly, to access a website, say www.google.com? Should I ph…

>When it comes to physical properties there's a huge difference between reading a banner posted in a street and entering the property to read some secret data: you have to be in different locations. That's why your analogy is completely faulty. At no point is accessing a web server similar in any matter to reading words off of a banner posted in a street. You cannot use a faulty analogy of your own to describe why my…

The "don't walk into someone else's house" rule applies to ALL houses everywhere. You are explicitly forbidden to enter a house unless explicitly authorized.

When it comes to website, there are billions of domains in the planet, each one has multiple internal URLs, ranging from tens to several million. You can't expect everyone to have common knowledge about every domain and link. It is beyond ridiculous to compare the two.

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#152

This is actually bad, would not it be better if sites would be allowed to block crawlers? I don't see what is the legal basis for forbidding to ban scrapers. Is there a law that a site must serve pages for anyone?

Yes - for instance if your e-commerce site banned people from visiting based on whether their zip code made it more likely they were of a certain racial group, you would be running afoul of the law.

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#153

Earlier quoted context omitted.

I would assume the data would still be covered by copyright meaning they could use that data and maybe create and sell derivative works, but not just scrape and publish.

My LinkedIn profile is copyright by me, insomuch as it’s a creative work.

The linkedin policy you agree to when you sign up indeed leaves you with ownership of content

"We will get your consent if we want to give others the right to publish your content beyond the Services."

But that license to publish would not extend to scrapers, so they would if re-publishing data, be violating copyright - it appears. Probably not for facts "Greg works at Widget Co." or statistics -- but probably for explicitly copying content, photos, etc. So linkedinclone.com (hypothetical) wouldn't be legal, but a service which scraped and reproduced facts would be.

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#154
post #74

Earlier quoted context omitted.

What does a CDN has to do with it?

Because if your application bundle is a fixed asset — like a JS SPA that fetches it’s data from an API then you can distribute your entire application via an inexpensive CDN. As soon as your application bundle is rendered on your servers dynamically then only part of your site can be delivered via CDN. Basically going all-JS gives you an app model where your sever side code doesn’t even know or care about HTML or the…

[deleted]

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#155
post #86
post #53

Earlier quoted context omitted.

> There is no reason why your page should refuse to load plain text without Javascript enabled. Sure there is. You prefer writing javascript and you want to serve your site through a CDN. You might not think that's a good reason, but that's certainly a reason.

Until the ADA comes along and demands you create an accessible to the blind site. I've often wondered when the laws would start to be applied and I think its coming

> It is a common misconception that people with disabilities don't have or 'do' JavaScript, and thus, that it's acceptable to have inaccessible scripted interfaces, so long as it is accessible with JavaScript disabled. A 2012 survey by WebAIM of screen reader users found that 98.6% of respondents had JavaScript enabled. [0]

and that was 7 years ago. There may be certain complicated interactions that are a challenge for screen readers but simply because a page relies on JavaScript for rendering doesn't automatically mean it is inaccessible to screen readers.

[0] https://webaim.org/techniques/javascript/#reliance

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#156
post #22

This action does more than that. The court left the preliminary injunction against LinkedIn in place: "The district court granted hiQ’s motion. It ordered LinkedIn to withdraw its cease-and-desist letter, to remove any existing technical barriers to hiQ’s access to public profiles, and to refrain from putting in place any legal or technical measures with the effect of blocking hiQ’s access to public profiles." So Lin…

Not allowing the CFAA to be (ab)used to attempt to make scraping illegal makes sense.

However, how is it reasonable to force a web site to serve its contents to a third-party company, without being allowed to make a decision whether to serve it or not? Serving the web site costs money, and the scraper surely isn't going to generate ad income...

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#157

This case is so ridiculous on multiple fronts that although this procedural ruling (injunction) seems technically correct (to allow the case to proceed to actual court), it could just as well have been thrown out with no difference in or ultimate harm to the parties. First, LinkedIn makes the claim that its users have a right to privacy against scraping by such a 3rd party. That's laughable. As the court saw, their w…

> suppose someone is taking your assets

Except that, in the digital sense, it's only copied. They now have it, but you didn't lose your assets or money besides the > So much craziness to go around.

I agree - I haven't read through the entire thing, but it looks like, instead of saying "you can't scrape", they could implicity give a license to users for personal and business use, but not be allowed the reselling of the data (of course carefully worded to allow the likes of Recruiters and whatnot to do so). It's like trying to argue that the DMCA says you can't create a torrent file of some movie.

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#158
post #22

This action does more than that. The court left the preliminary injunction against LinkedIn in place: "The district court granted hiQ’s motion. It ordered LinkedIn to withdraw its cease-and-desist letter, to remove any existing technical barriers to hiQ’s access to public profiles, and to refrain from putting in place any legal or technical measures with the effect of blocking hiQ’s access to public profiles." So Lin…

Does this prevent Google from returning captchas if you use a robot to scrape the search result pages, as they currently do?

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#159
how about websites using nonsense css-classes usually autogenerated through frameworks that make scrapping difficult ? I'm sure this ruling doesn't cover that case ? well globally I wish authorities would rule that public data should be published in computer accessible format e.g pdf's for humans and xblr's / csv for machines e.g in financial reports. lots of data in pdf's that costs a ton to mine. & tools like AWS Textract are hardly up to task.

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#160

> LinkedIn has taken steps to protect the data on its website from what it perceives as misuse or misappropriation. The instructions in LinkedIn’s “robots.txt” file—a text file used by website owners to communicate with search engine crawlers and other web robots—prohibit access to LinkedIn servers via automated bots, except that certain entities, like the Google search engine, have express permission from LinkedIn f…

Even if the case was tried today, 9th Cir. isn't binding on other regions of the US, and there's a bit of a split, as detailed in the opinion[1]:

> In recognizing that the CFAA is best understood as an anti-intrusion statute and not as a “misappropriation statute,” we rejected the contract-based interpretation of the CFAA’s “without authorization” provision adopted by some of our sister circuits. Compare Facebook, Inc. v. Power Ventures, Inc., 844 F.3d 1058, 1067 (9th Cir. 2016), cert. denied, 138 S. Ct. 313 (2017) (“[A] violation of the terms of use of a website—without more— cannot establish liability under the CFAA.”); Nosal I, 676 F.3d at 862 (“We remain unpersuaded by the decisions of our sister circuits that interpret the CFAA broadly to cover violations of corporate computer use restrictions or violations of a duty of loyalty.”), with EF Cultural Travel BV v. Explorica, Inc., 274 F.3d 577, 583–84 (1st Cir. 2001) (holding that violations of a confidentiality agreement or other contractual restraints could give rise to a claim for unauthorized access under the CFAA); United States v. Rodriguez, 628 F.3d 1258, 1263 (11th Cir. 2010) (holding that a defendant “exceeds authorized access” when violating policies governing authorized use of databases).

weev was tried in an area under the 3rd Cir. jurisdiction. Somewhat interestingly, his conviction was thrown out in 2014 on venue grounds (e.g. being tried in NJ), without addressing the statutory question.[2]

[1]: pp. 27-28 [2]: https://en.wikipedia.org/wiki/Weev?oldid=912921723#cite_ref-...

Post reply on HN