Live data from Hacker News

Congrats! Web scraping is legal! (US precedent)

parsers.me

301–310 of 409 posts

Re: Congrats! Web scraping is legal! (US precedent)

#301

Earlier quoted context omitted.

> but if the user does not have a service account (as is the case for HiQ, it doesn't seem they were using accounts for it), then your ToS does not apply, since you've technically not entered a binding legal contract with them. Are you sure about this? I am not a lawyer, but I believe that the Terms of Service applies to all users, not just those that explicitly set up a user account. I have interpreted the LinkedIn…

> Are you sure about this? I am not a lawyer, but I believe that the Terms of Service applies to all users, not just those that explicitly set up a user account. How would that even work? If I browse to any random public page of your website, it's served to me before you've even transmitted the terms of service. How could I be bound by those terms of service when I haven't even seen them?

As an engineer, I agree with what you are saying, but I think normal people and the courts disagree.

I think these sorts of contracts are called Adhesion Contracts (https://www.investopedia.com/terms/a/adhesion-contract.asp) and we interact with them all the time. For example, if you valet your car, the valet will hand you a piece of paper with a number printed on it to retrieve your car. On that paper you will find an adhesion contract that is valid and real (although not as powerful as the types of contracts that you sign)

Re: Congrats! Web scraping is legal! (US precedent)

#303
post #218

Earlier quoted context omitted.

> Would you be OK with a company scraping your blog and selling it? Selling it how? If they put my blog posts in a book and try to sell that book, that’s copyright infringement. If they put my blog posts in an ML model corpus to train a translation service, and they then charge pay-per-use access to the resulting service... I don’t think I’d care, nor do I think there’s anything morally or legally wrong with that. If…

Why can't I have terms on my website that say how you can use my information? Examples where this is allowed: - Images/media (Creative commons) - Code (Open source licenses) You say it isn't allowed for: - Personal data Unless I'm misunderstanding your philosophy (which seems to say copyright is OK, but public information must be public to all): You believe that it's morally OK for me to prevent a company selling my…

>Why can't I have terms on my website that say how you can use my information?

>- Personal data

So there's a couple of things in play here. You can't (generally) copyright facts - "Cthalupa is a Rocket Surgeon for the Space Force since 2001", if true, would not be something that I could get a copyright on.

https://www.newmediarights.org/business_models/artist/are_fa...

The second thing is that terms have to be agreed upon by both parties. If you give me information without us coming to an agreement on terms, I can't be bound by them. If you just put a link to a TOS on your website and don't require people agree to it before giving them access to data on your website, we did not enter into a contractual agreement.

Re: Congrats! Web scraping is legal! (US precedent)

#304
post #198

Perhaps linked in should switch to flutter . This type of technology will make it harder to scrape .

No, they should switch to Windows Azure Forms[0].

[0]: Clean-room reverse-engineered server-side replacement for Flutter, written in WebAssembly, which compiles source files to an .EXE using a .NET-enabled WebAssembly precompiler.

Re: Congrats! Web scraping is legal! (US precedent)

#305
post #213
post #11

The toxicity towards web-scraping is really what makes me lose hope in the current web. People want their data to be public and all of the benefits that comes with public data but then they want to chose who gets to see it - it's a complete and utter paradox. This precedent doesn't really mean much but is definitely step in the right direction.

The issue here for some, if not many, is a matter of scale. It is one thing if an end-user, whom I am trying to service, comes to my site and gets my publicly available data. Maybe I monetize with ads, maybe not. It doesn't matter, that is the audience I am trying to service, regardless of size. But when you scrape it my load goes up dramatically. A load I have to pay for. It is analogous to the privacy debates going…

So you throttle your users. We have http status codes for "too many requests" and all scraper software comes with a delay setting by default. Everybody who does scraping is supposed to know that its rude to blast a thousand requests per second.

Re: Congrats! Web scraping is legal! (US precedent)

#306
post #203

Earlier quoted context omitted.

It is true in the current situation, though I would prefer that we ensure free data must be free. In that case buyers of data would be incentivized to pressure providers of free data to improve the data quality.

The data does remain free, as long as LinkedIn still provides it for free. The data without the noise is what you're paying for. The service of winnowing out what you care about from what you don't care about. Considering how big of an effort it is, and that the source from which it came is still available, why should the cleaned data be free? If I collect fallen trees from public land, chop it into usable firewood,…

Absolutely spot-on.

I'm thinking of processed GIS data. If you have ever tried using the various formats that are supplied by government sites, you know what a huge pain it is.

I'm happy to pay a reasonable price for an interpreted and bowdlerized version.

Re: Congrats! Web scraping is legal! (US precedent)

#307

"HiQ only takes information from public LinkedIn profiles. By definition, any member of the public has the right to access this information. Most importantly, the appeals court also upheld a lower court ruling that prohibits LinkedIn from interfering with hiQ’s web scraping of its site." Surely I'm not reading this correctly. This would seem to suggest that websites are not legally allowed to prevent bots from crawli…

>Are those [ToS] legally void now?

They were pretty much legally void even before this precedent was established. They are only valid when they don't violate any existing U.S. law. Any authority assumed beyond that is completely false.

>can the court force me to revert that change?

No.

Re: Congrats! Web scraping is legal! (US precedent)

#308

"HiQ only takes information from public LinkedIn profiles. By definition, any member of the public has the right to access this information. Most importantly, the appeals court also upheld a lower court ruling that prohibits LinkedIn from interfering with hiQ’s web scraping of its site." Surely I'm not reading this correctly. This would seem to suggest that websites are not legally allowed to prevent bots from crawli…

> If LinkedIn wanted to force users to sign in to view profile info Do they not already do this? Every link I've ever seen for LinkedIn has redirected me to sign up page rather than showing me the content.

They want search engines to index their profiles and provide organic search results links to their site, but then those same sites will require you to sign in when clicking a link to another public profile. You can search for that 2nd profile in Google and then view it without signing in, but not by clicking internal links. I've experienced this with Quora, LinkedIn, Instagram, FB and others. They want to have their cake and eat it too.

Re: Congrats! Web scraping is legal! (US precedent)

#309

Earlier quoted context omitted.

Is that how really circuit court rulings get applied? I always understood each ruling on the rungs up the ladder to the supreme court applied across the land until a final ruling was determined.

Across the land within their circuit, over matters within their jurisdiction. Elsewhere the ruling is merely advisory in nature. It gets tricky with nationwide actors though. Besides some specialized topics like patents and international trade, nationwide orders and injunctions are the sort-of exception, which are based on a courts local jurisdictional power over a non-local nationwide actor. I'm not sure what the ge…

> One of the usual requirements for the Supreme Court to even hear a case is that different circuits have conflicting rulings on a matter.

That's not a “usual requirement”, it's one of many factors that can weigh in favor of the Supreme Court exercising discretionary appellate jurisdiction (and it's one that weighs very heavily in favor of it, even when no other favorable factors are present, since federal law meaning the same thing everywhere is an important principle.)

Re: Congrats! Web scraping is legal! (US precedent)

#310

"HiQ only takes information from public LinkedIn profiles. By definition, any member of the public has the right to access this information. Most importantly, the appeals court also upheld a lower court ruling that prohibits LinkedIn from interfering with hiQ’s web scraping of its site." Surely I'm not reading this correctly. This would seem to suggest that websites are not legally allowed to prevent bots from crawli…

> This is almost weirder. If LinkedIn wanted to force users to sign in to view profile info, would they be not allowed to do that because some company had signed a contract that implicitly assumed access to that data? If someone writes a web scraper for my site, and I unknowingly change my site in a way that breaks that scraper, can a court force me to revert the change?

LinkedIn has long wanted to have their cake and eat it too - they advertised that data as being publicly accessible and allow google to index specific user pages but then attempt to restrict other bots from crawling it.

If you have private data behind a login there isn't an issue here - if you have public data but want some people to login before viewing it (or not be able to view it) then that's where this ruling comes up. So, this mostly hits sneaky SEO folks and dark UX patterns that rely on tempting someone with accessible data and then pulling the rug out from them at the last minute.

If your website places data outside of authentication then everyone should be able to see that data... I'm curious to see the specifics around

> Surely I'm not reading this correctly. This would seem to suggest that websites are not legally allowed to prevent bots from crawling their sites. Lots of sites have ToS preventing such things, are those legally void now? Are captchas on public pages illegal, even if you request the page 8000 times in a second?

though - DoS attacks are clearly illegal, but with this precedent there's going to be a lot of back and forth to see where the line between DoS and scraping falls... and I think that makes this precedent a lot weaker than the headline would have you believe. A company can still threaten to drag you through a lot of litigation by accusing you of malicious page requests, it'll take a few cases to define where that line needs to fall.

Post reply on HN