Live data from Hacker News

Congrats! Web scraping is legal! (US precedent)

parsers.me

261–270 of 409 posts

Re: Congrats! Web scraping is legal! (US precedent)

#261
post #11

The toxicity towards web-scraping is really what makes me lose hope in the current web. People want their data to be public and all of the benefits that comes with public data but then they want to chose who gets to see it - it's a complete and utter paradox. This precedent doesn't really mean much but is definitely step in the right direction.

> People want their data to be public and all of the benefits that comes with public data but then they want to chose who gets to see it - it's a complete and utter paradox. That's a complete misconception. Of course you can find manufacture inconsistent ideologies if you combine ideas from different people, but I think you'd have a difficult time finding one person who believes what you just described. What I want i…

> I don't believe organizations have rights, period

So a group of people, joining together in a common cause, don’t have rights as members of that group?

You are contradicting yourself. Organizations are simply groups of people with a shared cause. To deny rights to the organization, you necessarily have to deny personal rights. Example: John and Sam form an organization for comic book collecting. They have a secret meeting to agree on a price to offer a potential seller of a rare comic book. Since John and Sam, Inc. “Have no rights” the seller is allowed to sit in the room during their meeting. However this violates the right of Sam to freely associate with John and his right to not have to associate with the interloper. It also violates the privacy of San and John as individuals since their private conversation — even as members of their two-man organization are now being shared with anyone who wants to sit in since, in this scenario, their organization doesn’t have rights.

It’s absurd.

Re: Congrats! Web scraping is legal! (US precedent)

#262

Earlier quoted context omitted.

It seems absurd if the 'interference' only directly affects their own property. Like, if my neighbors start monetizing livestreaming my backyard, suddenly I can't put up a fence? Except worse because in actuality, this third-party contract is costing them money through server load and bandwidth.

Your analogy doesn't hold. Your backyard is private property. The data that LinkedIn publishes is intended for the public. That's why Google can index the pages and give you results from LinkedIn.

Does that mean that ia grocery store offers free samples, I can go in every day and take all the samples, and the grocery is not allowed to selectively prevent me access?

Re: Congrats! Web scraping is legal! (US precedent)

#263

Earlier quoted context omitted.

Your backyard can be a walled garden--this is about the public front of the property.

Exactly your backyard is of course yours. But you are not at liberty to use it to damage others. There's lots of rules about this. For example, opening a brothel on your own land is definitely not legal without considering how it affects the neighborhood.

When is it "damage", and when is it "declining to contribute" ?

Re: Congrats! Web scraping is legal! (US precedent)

#264

"HiQ only takes information from public LinkedIn profiles. By definition, any member of the public has the right to access this information. Most importantly, the appeals court also upheld a lower court ruling that prohibits LinkedIn from interfering with hiQ’s web scraping of its site." Surely I'm not reading this correctly. This would seem to suggest that websites are not legally allowed to prevent bots from crawli…

> "Most importantly, the appeals court also upheld a lower court ruling that prohibits LinkedIn from interfering with hiQ’s web scraping of its site."

How would this affect Cloudflare's "checking your browser" anti-DDoS protection screen, meant to block bot requests from accessing sites?

Re: Congrats! Web scraping is legal! (US precedent)

#265

Earlier quoted context omitted.

I'm going to assume you're asking in good faith and try to address the confusion here. The human does get rights, the organization doesn't. In some cases, believing that humans have rights and believing that organizations have rights might lead one to the same action. In those cases, I'd take the action. I wouldn't want to violate a human's rights out of some vindictive dislike of organizations: that's not the point.…

Let's say that individual humans have the right to keep secrets. Let's also say that they have the right to keep secrets with their associates, and to tell them to who they please. Now, doesn't that make it legal for a group of people to keep secrets about you ? What about selling them? I just don't see what doing away with the legal fiction of corporate personage would do about Facebook.

"Now, doesn't that make it legal for a group of people to keep secrets about you?"

A group of people sure, but corporations are not people.

Re: Congrats! Web scraping is legal! (US precedent)

#266
post #258

"HiQ only takes information from public LinkedIn profiles. By definition, any member of the public has the right to access this information. Most importantly, the appeals court also upheld a lower court ruling that prohibits LinkedIn from interfering with hiQ’s web scraping of its site." Surely I'm not reading this correctly. This would seem to suggest that websites are not legally allowed to prevent bots from crawli…

> hiQ argued That does not mean that hte court agreed. The judges said that CFAA doesn't apply. In other words, the judges said that LinkedIn couldn't use the US legal system to force HiQ to stop. Judges didn't say that LinkedIn was barred from using technical measures. The court did allow a preliminary injunction against LinkedIn, due to the possibility of "monopolies" (to be determined in Court later), pending reso…

Something about malicious interfering with a contract? That might prevent technical measures?

Re: Congrats! Web scraping is legal! (US precedent)

#267

"HiQ only takes information from public LinkedIn profiles. By definition, any member of the public has the right to access this information. Most importantly, the appeals court also upheld a lower court ruling that prohibits LinkedIn from interfering with hiQ’s web scraping of its site." Surely I'm not reading this correctly. This would seem to suggest that websites are not legally allowed to prevent bots from crawli…

Linkedin want their data to be scraped by bots so they have to keep it public, otherwise you wouldn't find peoples profile from Google. They just don't want bots from from their competitors like hiQ to scrape it.

To me, this is crucial. If it's public and available for google, it's public and available for everyone. If you want content to be private, then make it private and accept that you won't get search engine traffic. Otherwise, don't be surprised when your publicly accessible content is accessed by gasp the public.

Re: Congrats! Web scraping is legal! (US precedent)

#268

Let’s not pretend this is a pure win. There are good uses of web scraping, like Archive.org trying to preserve the web. But what HiQ is doing is looking at public LinkedIn profiles and then snitching to employers if they think an employee is searching for a new job. It’s easy to blanket say “web scraping is legal, do what you will“. The tricky part is protecting people’s public data while not giving a huge moat to gi…

That's the thing. Web scraping isn't really the problem here. It's what companies are doing with personal information. If LinkedIn started doing the same thing as HiQ, it would be just as bad (probably worse), but the legality of web scraping is irrelevant to that.

That's a good point, and we certainly should write our data protection laws to prevent LinkedIn from doing the same thing — but there's a crucial difference between the two.

I've consented to give my data to LinkedIn, and I can withdraw my consent and data if they start doing something I don't like. On the other hand, hiQ has vacuumed up my data without my consent, and there's really no way for me to stop them other than retreating from my public profile. Certainly the legality of web scraping is relevant to that.

Re: Congrats! Web scraping is legal! (US precedent)

#269
post #237

Earlier quoted context omitted.

Provide an API for public data to reduce the costs associated with rendering a full blown page, and deliver just the information needed.

Who pays for that API and the bandwidth? What’s in it for the data provider? On LinkedIn, viewing the data now shows ads or at least prompts the viewer to join the network. With scrapping and free API access, how exactly does LinkedIn benefit for their work of hosting the data?

My guess is hiQ (and others) would happily pay for an API over the data they're scraping right now.
Post reply on HN