Live data from Hacker News

9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

cdn.ca9.uscourts.gov

261–270 of 293 posts

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#261

Earlier quoted context omitted.

I mean it should. That is also a huge anti-competitive action that just isn't pursued by anyone yet: google makes money of scraping and denies scrapers scraping them - that's all sorts of messed up. The problem is that someone would have to sue google first and no one will do that unless there's big business incentive and big business can already scrape the shit out of google. This is the weird thing about web-scrapi…

Google only scrapes sites that allow it by their robots.txt file so I don’t think their policy is as hypocritical as you are making it sound.

Google has an effective monopoly on search engine market - you _can't realistically_ block google from scraping your website. They have this power and they are abusing it.

Also robots.txt is bullshit, if a person can access a public website why automated script shouldn't - technically speaking it's the same thing.

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#262

Earlier quoted context omitted.

> However, how is it reasonable to force a web site to serve its contents to a third-party company, without being allowed to make a decision whether to serve it or not? Your statement makes absolutely no sense. That's not how internet works. If you serve something publicly you don't get to cherry pick who sees it. Not only it makes no sense technically it's also a huge anti-competitive case.

Of course you get to choose. You can reject requests based on their user agent, their IP address, the owner or likely geographic location of the IP address, and many other possibilities.

What are these possibilities? You only get IP and client side information that client is _willingly_ sending to you. So if a script/user/bot/etc tells it's Firefox from 1.2.3.4 then all you know that it's a request from 1.2.3.4 that says it's Firefox. You can ask it to run Javascript code but that's beyond classic web interaction and then again you need to trust the client.

This interaction is impossible to be trustless thus every client can only be served based on their IP or some convoluted, hack exchange that is cat-and-mouse game at best.

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#263
post #22

This action does more than that. The court left the preliminary injunction against LinkedIn in place: "The district court granted hiQ’s motion. It ordered LinkedIn to withdraw its cease-and-desist letter, to remove any existing technical barriers to hiQ’s access to public profiles, and to refrain from putting in place any legal or technical measures with the effect of blocking hiQ’s access to public profiles." So Lin…

Not allowing the CFAA to be (ab)used to attempt to make scraping illegal makes sense. However, how is it reasonable to force a web site to serve its contents to a third-party company, without being allowed to make a decision whether to serve it or not? Serving the web site costs money, and the scraper surely isn't going to generate ad income...

LinkedIn’s public facing content is exactly that: public. This ruling merely says accessing public content isn’t hacking and so LinkedIn cannot use the CFAA as discriminatory weapon to limit access to that public facing content.

If LinkedIn wants to block access they need to do so by another means that isn’t described as hacking.

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#264
post #246

Earlier quoted context omitted.

I would argue that under spirit of net neutrality you either serve your site to everyone equally(the public facing part) or to no one. Hosting costs money, servers cost money.. but maybe create a public facing API that is way cheaper and easier to use than scraping your website? I see that ruling in positive light that it might promote more open and structured access to the public facing data.

Why should you be forced to serve content to people who won't look at your ads?

Like disabled users with screen-readers?

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#265

Earlier quoted context omitted.

I never said that fact can be copyrighted, I said that most of the things people put around in their profile can be. I was responding to the claim that the data were not under copyright made above. If you just scrap name, company, position, this is fine, but I highly doubt that they just do that. This lawsuit can have tons of side effects.

I think what hiQ does is to predict whether a particular employee is about to quit. So the interesting question to me is whether you can lawfully make predictions based on published information if that information is under copyright. In Europe the answer is probably no, because the assumption is that in order to analyse data you have to copy it first. To me, this interpretation of the term "copying" makes very little…

Europe has database rights, which has a fair dealing exemption for data analysis.

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#266
post #22

This action does more than that. The court left the preliminary injunction against LinkedIn in place: "The district court granted hiQ’s motion. It ordered LinkedIn to withdraw its cease-and-desist letter, to remove any existing technical barriers to hiQ’s access to public profiles, and to refrain from putting in place any legal or technical measures with the effect of blocking hiQ’s access to public profiles." So Lin…

Not allowing the CFAA to be (ab)used to attempt to make scraping illegal makes sense. However, how is it reasonable to force a web site to serve its contents to a third-party company, without being allowed to make a decision whether to serve it or not? Serving the web site costs money, and the scraper surely isn't going to generate ad income...

I think the title is wrong and they are holding that linkedin can not block specifically hiq from viewing public data.

Which seems fair its public or its not, you can't pick and choose who its public for and who is a second class citizen.

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#267
post #22

This action does more than that. The court left the preliminary injunction against LinkedIn in place: "The district court granted hiQ’s motion. It ordered LinkedIn to withdraw its cease-and-desist letter, to remove any existing technical barriers to hiQ’s access to public profiles, and to refrain from putting in place any legal or technical measures with the effect of blocking hiQ’s access to public profiles." So Lin…

Not allowing the CFAA to be (ab)used to attempt to make scraping illegal makes sense. However, how is it reasonable to force a web site to serve its contents to a third-party company, without being allowed to make a decision whether to serve it or not? Serving the web site costs money, and the scraper surely isn't going to generate ad income...

It's not forcing anything. Don't make a page public then? If a page is public then it is fair game.

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#268
post #25

A choice quote: > In recognizing that the CFAA is best understood as an anti-intrusion statute and not as a “misappropriation statute,” Nosal I, 676 F.3d at 857–58, we rejected the contract-based interpretation of the CFAA’s “without authorization” provision adopted by some of our sister circuits. Compare Facebook, Inc. v. Power Ventures, Inc., 844 F.3d 1058, 1067 (9th Cir. 2016), cert. denied, 138 S. Ct. 313 (2017)…

The post is now here:

https://reason.com/2019/09/09/scraping-a-public-website-does...

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#269

Earlier quoted context omitted.

> under spirit of net neutrality you either serve your site to everyone equally(the public facing part) or to no one Huh? Net neutrality isn't about the server or client... it's about the network operator in between them.

I suspect Xelbair is making a more expensive definition of net neutrality, taking as a basis the one that says it's about network operators only.

I think you wanted to say more expansive? But it's definitely also more expensive. :D

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#270

Earlier quoted context omitted.

I think what hiQ does is to predict whether a particular employee is about to quit. So the interesting question to me is whether you can lawfully make predictions based on published information if that information is under copyright. In Europe the answer is probably no, because the assumption is that in order to analyse data you have to copy it first. To me, this interpretation of the term "copying" makes very little…

Europe has database rights, which has a fair dealing exemption for data analysis.

I'm not sure what "database rights" refers to specifically, but the whole matter is actually rather complicated, because the EU copyright directive has a lot of optional exceptions that member states may or may not adopt.

Most of these exceptions only apply to non-commercial use though. So they wouldn't apply in a case like hiQ.

UK specific exceptions are explained here:

https://www.gov.uk/guidance/exceptions-to-copyright

Unfortunately, both Labour and the Tories have taken a relatively hard line in the EU copyright negotiations, so it seems unlikely that things will be relaxed very much after Brexit.

Post reply on HN