Live data from Hacker News

Congrats! Web scraping is legal! (US precedent)

parsers.me

101–110 of 409 posts

Re: Congrats! Web scraping is legal! (US precedent)

#101
post #11

The toxicity towards web-scraping is really what makes me lose hope in the current web. People want their data to be public and all of the benefits that comes with public data but then they want to chose who gets to see it - it's a complete and utter paradox. This precedent doesn't really mean much but is definitely step in the right direction.

One of my clients is involved in property tax collection and reporting. Property Tax records are public info, and their website allows looking up the records for any property without a login. However, the data behind this website it the _source_ of the public records, and not the public records themselves (which would be local government databases). For years now we've been in an arms race with someone using a botnet…

I'm in a similar job. We block people from scraping if they break a threshold, but we also refer them to the reporting system, which can get all of the information that they are collecting in a variety of formats.

I wonder if something like this would be allowed: if all the public information was available in a well-collated format, then can scrapers be blocked? I imagine that will eventually be fought in court as well.

Re: Congrats! Web scraping is legal! (US precedent)

#102

"HiQ only takes information from public LinkedIn profiles. By definition, any member of the public has the right to access this information. Most importantly, the appeals court also upheld a lower court ruling that prohibits LinkedIn from interfering with hiQ’s web scraping of its site." Surely I'm not reading this correctly. This would seem to suggest that websites are not legally allowed to prevent bots from crawli…

The court document says "... refrain from putting in place any legal or technical measures with the effect of blocking hiQ's access to public profiles." on page 11. I wonder if they mean targeted measures specifically blocking hiQ but allowing others such as Google.

Re: Congrats! Web scraping is legal! (US precedent)

#103
post #26
post #11

The toxicity towards web-scraping is really what makes me lose hope in the current web. People want their data to be public and all of the benefits that comes with public data but then they want to chose who gets to see it - it's a complete and utter paradox. This precedent doesn't really mean much but is definitely step in the right direction.

> People want their data to be public People don't want their data to be public. People want other people's data to be public. One's own data everyone thinks should be private and tightly controlled. This applies to people and businesses equally.

I think the GP in this context means, e.g. LinkedIn wants their information to be public in the cases where it benefits them as a business. But then they want it to not be public when it doesn't benefit them. There is no such thing as "public information, except ..." - information is public, or it's not. If none of LinkedIn's data was public, they would have a much harder time getting people to sign up, and having as many users signed up as possible is part of their business model.

From a copyright perspective (since that's what LinkedIn's lawyers claimed): imagine if a newspaper sued another newspaper, saying that - not just the content of its paper - the information in the newspaper was copyrighted and could not be accessed by "unauthorized" third party companies. Either you print it, or you don't!

Re: Congrats! Web scraping is legal! (US precedent)

#104
post #21

Earlier quoted context omitted.

Why would a review be a copyrightable creative work, while a LinkedIn resume wouldn't be?

I think perhaps the layout, cover letter, and maybe any flourishing notes are copyrightable, but the actual details of work experience and education are not.

Yeah, I would think the "description" section for each job would be copyrightable, but the simple "title", "company", "year" fields would not be.

Re: Congrats! Web scraping is legal! (US precedent)

#106
post #50

Earlier quoted context omitted.

There’s a very effective way out, don’t your data on the public web.

[decided to delete because I misunderstood the context]

Let's not conflate a) a person's personal data, and b) a business's dataset. The GP and the article are clearly referring to the latter. Preventing web scraping won't protect users from businesses collecting their data.

Re: Congrats! Web scraping is legal! (US precedent)

#107

Let’s not pretend this is a pure win. There are good uses of web scraping, like Archive.org trying to preserve the web. But what HiQ is doing is looking at public LinkedIn profiles and then snitching to employers if they think an employee is searching for a new job. It’s easy to blanket say “web scraping is legal, do what you will“. The tricky part is protecting people’s public data while not giving a huge moat to gi…

That's the thing. Web scraping isn't really the problem here. It's what companies are doing with personal information. If LinkedIn started doing the same thing as HiQ, it would be just as bad (probably worse), but the legality of web scraping is irrelevant to that.

Re: Congrats! Web scraping is legal! (US precedent)

#108
post #17

Earlier quoted context omitted.

Yes you can scrape them, no you cannot repubilsh them. Everything you listed is protected by copyright. You cannot infringe on copyrights because of this ruling. >hiQ argued that LinkedIn’s technical measures to block web scraping interfere with hiQ’s contracts with its own customers who rely on this data. In legal jargon, this is called” malicious interference with a contract”, which is prohibited by American law Do…

I think any ruling that says LinkedIn can't put in protectionary measures against automated requests is doomed to be overturned, as long as they're not doing it discriminately. Captcha, rate limiting, user agent testing, etc are all common tools to protect against malicious/unintentional denials of service. The question is what was LinkedIn doing, and did it specifically target hiQ while permitting others of the same…

Why would it be an issue if it is discriminatory? Linkedin can use its servers any way they like, unless they ve promised their users that their data can be scraped indiscriminately

Re: Congrats! Web scraping is legal! (US precedent)

#109
post #71

Earlier quoted context omitted.

I build https://awardfares.com together with a friend which scrapes airlines' award seat availability. Airlines' websites are horrible from a UX perspective so scraping the data and presenting it in a better way was a pretty obvious use case.

No British Airways? A quick look shows a lot of award availability to China for some reason...

It's coming in the next couple of weeks! Both me and my co-founder had Star Alliance frequent flyer programmes so that's what we focused on first.

Re: Congrats! Web scraping is legal! (US precedent)

#110
post #11

The toxicity towards web-scraping is really what makes me lose hope in the current web. People want their data to be public and all of the benefits that comes with public data but then they want to chose who gets to see it - it's a complete and utter paradox. This precedent doesn't really mean much but is definitely step in the right direction.

I want web scraping to be legal—but, is it really contradictory to say "I want this data to be accessible to real humans only"? Any person can post on Hacker News. However, if someone made a bot to post to Hacker News, I think most of us would be pretty upset.

Scraping information is not the same as posting. There are a number of bots that scrape Hacker News and people here generally consider them pretty cool.
Post reply on HN