Live data from Hacker News

Congrats! Web scraping is legal! (US precedent)

parsers.me

191–200 of 409 posts

Re: Congrats! Web scraping is legal! (US precedent)

#191
post #119

Earlier quoted context omitted.

> Surely I'm not reading this correctly. This would seem to suggest that websites are not legally allowed to prevent bots from crawling their sites. Lots of sites have ToS preventing such things, are those legally void now? Are captchas on public pages illegal, even if you request the page 8000 times in a second? This is just a preliminary injunction. This wasn't an actual ruling on the case. This just says that unti…

You don’t understand what a preliminary injunction is then. It’s a very, very strong indication that they will win. Courts don’t issue preliminary injunctions unless it’s extremely likely the side who won the preliminary injunction will win.

It only requires a “substantial” likelihood that side will win (not an “extreme” one), which basically means there’s a substantive dispute. The more difficult criterion is a substantial likelihood that irreparable harm will occur if the injunction isn’t granted (irreparable harm is supposed to be a pretty extreme thing — it means you can’t fix it with any amount of money).

Re: Congrats! Web scraping is legal! (US precedent)

#192
post #134

Since this is probably granted for anyone with a technical understanding, it is nice to see that the legislative powers are on board with this.

Just because something is possible technically doesn't make it ok legally. I think there's still various issues though, as per GPDR I don't think another company can just copy that data from Linkedin. That it's easily visible doesn't matter for GDPR.

But the offender would be LinkedIn if it exposed personal data, even per GDPR.

The cases were people were convicted by getting certain webpages is addressed here.

Re: Congrats! Web scraping is legal! (US precedent)

#193
post #11

The toxicity towards web-scraping is really what makes me lose hope in the current web. People want their data to be public and all of the benefits that comes with public data but then they want to chose who gets to see it - it's a complete and utter paradox. This precedent doesn't really mean much but is definitely step in the right direction.

I want web scraping to be legal—but, is it really contradictory to say "I want this data to be accessible to real humans only"? Any person can post on Hacker News. However, if someone made a bot to post to Hacker News, I think most of us would be pretty upset.

What you might mean is:

- I want my data to be publicly available

- I don't want my data to be processed/distributed/sold without my permission

E.g. individual use is fine, profit making is not.

Which is my expectation with LinkedIn. I want people to see my profile, I don't want them to sell it as marketing leads!

Re: Congrats! Web scraping is legal! (US precedent)

#195

Earlier quoted context omitted.

> Lots of sites have ToS preventing such things, are those legally void now? Are captchas on public pages illegal, even if you request the page 8000 times in a second? ToS are subservient to the law; you can (probably) terminate a service account from a user that breaks your ToS, but if the user does not have a service account (as is the case for HiQ, it doesn't seem they were using accounts for it), then your ToS do…

There's a long, long history (probably hundreds –if not thousands– of years old) of selling aggregated or processed publicly-available information. I'm not particularly thrilled with it, but enough people think of it as a valuable enough service to pay for; even if they know they could get it themselves, for free. LinkedIn users (as opposed to the company) might actually like what HiQ is doing, as it may help their o…

> but enough people think of it as a valuable enough service to pay for; even if they know they could get it themselves, for free.

It's not free, it takes time to collect data. Buying it makes a lot of sense as long as you pay less than what's your own time worth to you...

Re: Congrats! Web scraping is legal! (US precedent)

#196

Earlier quoted context omitted.

https://en.wikipedia.org/wiki/Tortious_interference This would mostly mean that you cannot start interfering with webscraping you previously allowed merely because you learned that they're making money with the scraped data.

It seems absurd if the 'interference' only directly affects their own property. Like, if my neighbors start monetizing livestreaming my backyard, suddenly I can't put up a fence? Except worse because in actuality, this third-party contract is costing them money through server load and bandwidth.

User data is not theirs property.

Re: Congrats! Web scraping is legal! (US precedent)

#197
post #11

The toxicity towards web-scraping is really what makes me lose hope in the current web. People want their data to be public and all of the benefits that comes with public data but then they want to chose who gets to see it - it's a complete and utter paradox. This precedent doesn't really mean much but is definitely step in the right direction.

I want to be able to use LinkedIn to network with colleagues and people in my industry. If someone wants to scrape my profile to make a report on industry trends, I’m fine with it. What I don’t want is hiQ vacuuming up my data so they can snitch to my employer if they think I’m job hunting. How is this a paradox? Tech — the web in particular — is supposed to be an equalizing force, but HiQ is clearly trying to give m…

The thing is, GDPR has theoretically solved this in the EU. The UK's ICO is about to publish guidance prohibiting scraping public user information for marketing (where the user would not expect it to be used for that).

It's a really easy solution, because companies need to prove how they got your data when asked.

When you track the source of the mailing list you're getting spam from and they say "We scraped it from LinkedIn", they get fined.

Re: Congrats! Web scraping is legal! (US precedent)

#199
I'm seeing some confusion about how this affects people outside of the West Coasts 9th circuit. It certainly affects you if you are in other circuits. There's no question that this precedent will be brought up in other courts, even if they aren't bound to uphold the decision.

IANAL, but a bit of a supreme court hobbyist. For those unfamiliar, a short lesson on how the federal court system works that I think would be useful for HN readers:

Generally speaking, a case first goes to District Court. There are 94 districts in the country and they tend to consist of small regions. For example, Northern California (note these are for federal courts... each state has it's own state system, which operates differently in each state). So you first go to one of these courts when you bring a lawsuit.

Now, the lawsuit is decided and one side is unhappy with the verdict. In a lot of cases, they might say, "ok, I'm unhappy, but I'm also spending a lot of money on lawyers, so i'll accept the decision and move on". Or they have the right to appeal the decision. If they choose to do the latter, then the next layer of the court system, the Circuit Courts come into play.

The Circuit Courts consist of 11 "circuits" plus the DC [0] and Federal Circuits [1]. When you appeal a case to one of the circuit courts, they have to hear your case. It's your right to appeal. However, if they think the case doesn't have merit, they can issue what's called a "summary judgement" where they issue an opinion without a full trial. In other words, if they think the lower court issued the correct decision and they think it would be a waste of time to go through a trial to appeal it, they can look over the facts of the case and make a decision without a trial.

At the next level up, you have the Supreme Court [2]. Unlike at the circuit court level, you have no right to appeal to the supreme court. Generally speaking, they get to decide which cases they want to take, so if you think that the appeals court screwed up and the supreme court decides to not take your case, there's nothing you can do. Unlike appeals courts, they aren't even obligated to look over the facts of the case at all if they don't want to.

Instead, what happens is that you petition the Supreme Court to hear your case. So you lose your case at appeals court and you basically file paperwork with the SC saying "please please hear my case, here's why i think you should".

The SC turns down a lot more cases than it hears. So what makes them take a case on? One is if the case is super super important. Something like a case against Obamacare or something else of very high national significance. But typically, the SC takes a lot of boring cases too, and for the most part, this has to do with circuit splits. A circuit split is when two of the regional circuit courts issue conflicting rulings. So for example, if every circuit more or less rules the same way on a given issue, there's nothing for the supreme court to decide on. The system is working as intended. But if two circuits disagree, then the SC's job is to resolve the issue so that federal law is applied uniformly.

So in this case, the ruling was issued in the Ninth Circuit (which covers California and much of the West). Technically speaking, I can sue someone over scraping my site in Wisconsin (in the 7th Circuit) and the judge can rule in my favor (that scraping is illegal) since she isn't bound to follow 9th Circuit appeals rulings, whereas a district judge in Colorado or Montana (both in 9th Circuit) are bound by the appeals court precedent. And then they appeal the case in Chicago and the appeals court also isn't bound by the Ninth Circuit ruling. But you can be damn sure that the lawyers for the defense are going to bring up that Ninth Circuit case as precedent. And generally speaking, the precedent does matter (circuits don't want to create circuit splits).

Coming back to this case, does this mean that you have carte blanche to scrape websites anywhere in the country? No, case law is going to need to evolve more to get to the point where you can safely think that way. But, this is definitely an important step in that direction.

[0] Washington DC has its own circuit, even though it's just a city, not a region. This seems odd at first glance, but in fact, a lot of lawsuits against the federal government come through this circuit, which is why DC gets it's own circuit, while for example, NYC does not.

[1] The United States Court of Appeals for the Federal Circuit is a special case. Most of these circuit courts have to do with regions. So if I commit a federal crime in Florida, it gets tried in a Florida District court and then it gets appealed in the seat of the 11th circuit, which is in Atlanta. However, certain cases, based on subject material, don't get appealed in Atlanta, but go instead to the Federal Circuit. These tend to be things with national ramifications. Patents are a prime example, where you really really don't want a patent being enforced in Iowa, but not in Alabama. So for these special cases, we set up a different appeals system.

[2] There are a few times when cases go straight to the supreme court. For example, in disputes between two states or cases involving ambassadors and other public ministers, a case might go straight to the SC and skip the lower courts. But this is the exception, not the rule.

Re: Congrats! Web scraping is legal! (US precedent)

#200

Earlier quoted context omitted.

One of my clients is involved in property tax collection and reporting. Property Tax records are public info, and their website allows looking up the records for any property without a login. However, the data behind this website it the _source_ of the public records, and not the public records themselves (which would be local government databases). For years now we've been in an arms race with someone using a botnet…

Wouldn't the solution be to offer a streamlined download (maybe even as a torrent if you're worried about bandwidth) of all the data then?

If the scraper contacted the client, said what they need the data for, and (probably) paid for api access, then my client would probably go for it.

My client is under no obligation to make access to this data easier. It's not really their data either; the information is property addesses, owner names and addresses, and tax assessments and payments. My client wouldn't want to make it easier for scammers to get that data. So they're not going to do anything unless they know the scraper is legit. If that's the case, the api would require authentication, and any fees would be for the server load, not the data.

Post reply on HN