Live data from Hacker News

9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

cdn.ca9.uscourts.gov

231–240 of 293 posts

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#231

Earlier quoted context omitted.

Facts can't be copyrighted, so such things as whether or not a person worked for a certain company, or went to a certain school, are unprotected, and with this ruling can be scraped, at least in the U.S. Others things common on LinkedIn, as you rightly point out, are protected--but by copyright law, not the CFAA. So a scraper acting in good faith would have to be careful about what they used if they wanted to respect…

I never said that fact can be copyrighted, I said that most of the things people put around in their profile can be. I was responding to the claim that the data were not under copyright made above. If you just scrap name, company, position, this is fine, but I highly doubt that they just do that. This lawsuit can have tons of side effects.

I think what hiQ does is to predict whether a particular employee is about to quit.

So the interesting question to me is whether you can lawfully make predictions based on published information if that information is under copyright.

In Europe the answer is probably no, because the assumption is that in order to analyse data you have to copy it first.

To me, this interpretation of the term "copying" makes very little sense. So I wonder what US law makes of it.

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#232

Earlier quoted context omitted.

As a real-world analogue: you can indeed be guilty of trespassing on someone's property even if you don't have to jump over any fences or pick any locks to get there. In some places, they don't even have to have a "no trespassing" sign. Simply being present on someone else's property without an invitation from them is illegal, and no, an open door does not count as an invitation.

this is bad comparison because when scraping a site, you don't cross any borders, you just send and receive information. You can compare this to a phone call or to talking to someone.

Those are apple to orange comparisons: A phone call - you don't have to answer the call, nor say anything once you know (or don't know) who the caller is or what their intention is - you can stop whenever you want - is it a robo-call? You hangup. And similarly with talking to someone (in person - if they say something you're free to just not respond; and if they persist, it's harassment.

The main reason I see businesses being concerned about being required to serve scrapers pages (even at a reasonable rate of download) is that there's still cost associated to it, and more so the more scrapers try to access and regularly access the data for updates. Similarly, if it is the users of a platform who have input the data, update it, and they are only wanting it presented on that platform (for whatever reasons) then what rights do they have?

Is the answer then requiring adding another acknowledgement message like "this site uses cookies" required, perhaps with required response before moving forward to have users acknowledge "scraping isn't allowed" - akin to "no trespassing" signs on properties? That seems awfully ridiculous to put the onus on 100% of users (including the overwhelming majority being non-scrapers), adding friction and speed of access to billions of internet surfers? Of course browsers could then could act as a layer that auto-respond to that or pre-agree to the rules - perhaps in a way reading through a site's TOS and pre-approving what you agree to. And as the trend has been otherwise it leads to closed platforms so the data isn't considered public; I won't argue whether that is good or bad for the general internet, however how much value is there in a person having access to that data without having to be a user?

Or the much simpler thing is we could put the onus on businesses who are scraping or will use scraped data to not cause this mass friction.

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#234
post #22

This action does more than that. The court left the preliminary injunction against LinkedIn in place: "The district court granted hiQ’s motion. It ordered LinkedIn to withdraw its cease-and-desist letter, to remove any existing technical barriers to hiQ’s access to public profiles, and to refrain from putting in place any legal or technical measures with the effect of blocking hiQ’s access to public profiles." So Lin…

And I presume it would also apply to cloudflare’s captchas.

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#235
post #22

This action does more than that. The court left the preliminary injunction against LinkedIn in place: "The district court granted hiQ’s motion. It ordered LinkedIn to withdraw its cease-and-desist letter, to remove any existing technical barriers to hiQ’s access to public profiles, and to refrain from putting in place any legal or technical measures with the effect of blocking hiQ’s access to public profiles." So Lin…

Not allowing the CFAA to be (ab)used to attempt to make scraping illegal makes sense. However, how is it reasonable to force a web site to serve its contents to a third-party company, without being allowed to make a decision whether to serve it or not? Serving the web site costs money, and the scraper surely isn't going to generate ad income...

Surely the action is "if you display stuff in public you can't segment the public".

You're not obliged to have public access.

Is there perhaps a factor here of users having an expectation that their profile is publicly accessible; so companies hosting that profile shouldn't be able to choose _secretly_ "who" can access it?

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#236
post #230
post #148

Earlier quoted context omitted.

> Plus Google only exists from scraping content, but I believe their TOS includes "don't scrape our content". Yes. This is EXTREMELY frustrating. Of all companies to prevent scraping, Google is the most ironic. Especially since their goal is to organize the world's information, it shocks me that there's no way to get access to this organized information from machine to machine.

Perhaps this issue will be recognised in some of the antitrust investigations. If I am not mistaken, they no longer claim "organize the world's information" as their goal.

But it still is:

> https://about.google/

"Our mission is to organize the world’s information and make it universally accessible and useful."

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#237
post #189
post #163

Earlier quoted context omitted.

Isn't the issue of being selective on who can view the content? If I, random Joe User views the publicly available content you have no issue. But if someone scrapes that data them you'd want to charge them. Unless I click on the ad, the act of using your bandwidth doesn't change based on who the viewer is. You'd want to apply fees based on the future use of the data rather than on your actual costs.

I'd assume if you weren't signing up, you'd probably look at like 10 profiles tops. A scraper is more than likely going to run through anything and everything it can grab links to (provided it doesn't leverage a very specific filtering mechanism for selecting profiles to scrape). I could see the hit from a scraper being heavier than that of a typical user. There's also the potential that a user is going to click an a…

Presumably you'd be allowed to limit a scraper to a standard user bandwidth, and a standard user access - X links per day, Y bandwidth.

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#238
post #230

Earlier quoted context omitted.

Perhaps this issue will be recognised in some of the antitrust investigations. If I am not mistaken, they no longer claim "organize the world's information" as their goal.

But it still is: > https://about.google/ "Our mission is to organize the world’s information and make it universally accessible and useful."

Appears I am mistaken. Cheers.

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#239
post #210
post #196

Earlier quoted context omitted.

From reading the opinion, I think the argument goes something like this: > First, LinkedIn does not contest hiQ’s evidence that contracts exist between hiQ and some customers, including eBay, Capital One, and GoDaddy > Second, hiQ will likely be able to establish that LinkedIn knew of hiQ’s scraping activity and products for some time. LinkedIn began sending representatives to hiQ’s Elevate conferences in October 201…

That’s quite ... crazy. Be restaurant. Be on Deliveroo. Be getting low margins because of high fees. So basically you can’t decide not to use Deliveroo any more, to improve margina (“secure an exonomic advantage”). I mean, you can cancel Deliveroo, but only as long as you’re not “inducing a breach of their contract”. So only a matter of time before Deliveroo writes a contract “we’re obligated to deliver food for you…

Choosing not to use a middleman any more so that you can secure higher margins sounds like about clearest example of a "legitimate business reason" imaginable. The purpose of the act is to immediately increase your margins, not to hurt Deliveroo because you don't want their competition.

That's very different from the case in question, where LinkedIn's motive for cutting off hiQ's access is to inflict damage on hiQ because they are a potential competitor.

Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]

#240
post #117
post #73

Earlier quoted context omitted.

It's limited to public pages. They can still discriminate whom they serve, with logins or something, but they can't limit your ability to access their page in a way that you prefer.

Of course they can. The question is whether they can do so legally and, if not, based on what laws.

[deleted]
Post reply on HN