Live data from Hacker News

Congrats! Web scraping is legal! (US precedent)

parsers.me

181–190 of 409 posts

Re: Congrats! Web scraping is legal! (US precedent)

#182
post #36

German copyright has the concept of a "Datenbankwerk" (since the 90s). E.g. the telephone book contains lots of boring facts that are each in themselves not copyrightable. However the collection in itself is copyrightable, as it required substantial effort to create. It seems odd that US copyright law wouldn't have a similar provision, or that it doesn't apply here?

US law has a very similar approach. See e.g. https://www.bitlaw.com/copyright/database.html

But I don't think that's relevant here, as a) this isn't a copyright case and b) HiQ are not attempting to recreate the entire "compilation" of LinkedIn

Re: Congrats! Web scraping is legal! (US precedent)

#183
post #119

Earlier quoted context omitted.

> Surely I'm not reading this correctly. This would seem to suggest that websites are not legally allowed to prevent bots from crawling their sites. Lots of sites have ToS preventing such things, are those legally void now? Are captchas on public pages illegal, even if you request the page 8000 times in a second? This is just a preliminary injunction. This wasn't an actual ruling on the case. This just says that unti…

You don’t understand what a preliminary injunction is then. It’s a very, very strong indication that they will win. Courts don’t issue preliminary injunctions unless it’s extremely likely the side who won the preliminary injunction will win.

Huh, I thought in USA they also did them to avoid an injunction having the effect of making the judgement irrelevant. So, where the case is not clear cut the injunction could prevent one party acting to 'kill' the other (and so avoid judgement) in the meantime?

Could you cite something on this that indicates this (my understanding here) is wrong?

Re: Congrats! Web scraping is legal! (US precedent)

#184
So many ideas start to come to mind if scraping is legal.

Can we start to scrape Google Search in order to bootstrap building an alternative to Google Search? Search is a really hard problem (that somebody should tackle), but if we can leverage what Google has already scraped from the web and associated with popular search terms, we can use that to help train and validate our search model.

Can we scrape Reddit, Twitter, or Facebook in order to stand up a competing service that strips out all the ads? It's hard to bootstrap a social media website, but if you can import all the content from the existing giants, your site is no longer a wasteland.

Can we finally scrape and get rid of IMDB? I'd love to put all of their content on a wiki and be done with it.

Re: Congrats! Web scraping is legal! (US precedent)

#185

Earlier quoted context omitted.

> This only affects the ninth circuit—which includes the tech hubs San Francisco, Seattle, LA, and Portland. It is only binding precedent in the Ninth Circuit, it is less accurate to say it only effects the Ninth Circuit, since decisions have effect other than as binding precedent.

Circuits can and do disagree.

Yes, that's why it is not binding elsewhere, which doesn't mean it has no effect. Particularly, if another Circuit has not issued a conflicting ruling, the Ninth Circuit ruling can be cited in and relied on by trial courts in that circuit as persuasive, rather than binding, precedent, so it can have an impact from the very earliest stages of the process.

A circuit split is also a reason for the Supreme Court to take a case, so the Ninth Circuit decision, without being binding, makes it less likely that any conflicting decision by another circuit will be the final resolution of the case in which that conflicting decision is issued, which is an important effect at the other end of the process.

Re: Congrats! Web scraping is legal! (US precedent)

#186
post #160

Earlier quoted context omitted.

https://en.wikipedia.org/wiki/Tortious_interference This would mostly mean that you cannot start interfering with webscraping you previously allowed merely because you learned that they're making money with the scraped data.

If I decide to change the class names, or HTML structure of the page, is that no longer allowed? How far does this go?

Are you doing it just to spite scrapers, i.e. with "malicious intent"? If you have some other reason, you won't be guilty of intentional tortious interference.

Re: Congrats! Web scraping is legal! (US precedent)

#187

Linkedin is taking this to the Supreme Court: https://www.law360.com/articles/1237505/linkedin-will-go-to-... No ultimate decision was ever made, and no, this doesn't make web scraping 100% legal. Wake me up when there's a new announcement because anyone interested in this already know this old news.

Bad look on LinkedIn, if you want the information to be visible publicly, don’t fault others for learning that information.

Gate the information behind a login then sue the scraper for violating TOS and not scraping itself, I can understand that.

Re: Congrats! Web scraping is legal! (US precedent)

#188
post #143
post #26

Earlier quoted context omitted.

> People want their data to be public People don't want their data to be public. People want other people's data to be public. One's own data everyone thinks should be private and tightly controlled. This applies to people and businesses equally.

In this case, LinkedIn users kind of do want their “public profiles” to be public. They’re online CVs; by definition, if you make one, your goal is to get it into the hands of anyone who asks for it! LinkedIn, likewise, has built its business model on an implicit contract with its users that it’s going to show their CV to anyone who asks for it. I think LinkedIn users would be surprised that LinkedIn doesn’t let bots…

>I think LinkedIn users would be surprised that LinkedIn doesn’t let bots read (scrape) their public CV.

LinkedIn is really really clear that:

- They won't share you information with 3rd parties

- You're not allowed to use information on LinkedIn for commercial purposes without their permission

- Other users can view your personal data

So, why would I expect random third party companies to be able to scrape and sell my personal information?

My personal information is there for the individual use of others, and for authorised use by recruiters (who are vetted/managed by LinkedIn).

I've chased down the convention spam mail I get using my GDPR rights, and surprise surprise, they got my details by scraping LinkedIn. That is absolutely not expected nor acceptable use of my data...

Re: Congrats! Web scraping is legal! (US precedent)

#189
post #11

The toxicity towards web-scraping is really what makes me lose hope in the current web. People want their data to be public and all of the benefits that comes with public data but then they want to chose who gets to see it - it's a complete and utter paradox. This precedent doesn't really mean much but is definitely step in the right direction.

I want to be able to use LinkedIn to network with colleagues and people in my industry. If someone wants to scrape my profile to make a report on industry trends, I’m fine with it. What I don’t want is hiQ vacuuming up my data so they can snitch to my employer if they think I’m job hunting. How is this a paradox? Tech — the web in particular — is supposed to be an equalizing force, but HiQ is clearly trying to give m…

We haven't solved that in the same way we haven't "solved" encryption not having a magical good people only door despite spook tantrums. There fundamentally isn't a possible mechanism and really wanting it doesn't change that.

It is a result of equality - not of outcome but rules. Open for everyone but those whose applications you don't like isn't open. On a technical level trying to prevent it is like the "evil bit" as a solution to malware.

Post reply on HN