Live data from Hacker News

Congrats! Web scraping is legal! (US precedent)

parsers.me

151–160 of 409 posts

Re: Congrats! Web scraping is legal! (US precedent)

#151
post #11

The toxicity towards web-scraping is really what makes me lose hope in the current web. People want their data to be public and all of the benefits that comes with public data but then they want to chose who gets to see it - it's a complete and utter paradox. This precedent doesn't really mean much but is definitely step in the right direction.

I want to be able to use LinkedIn to network with colleagues and people in my industry. If someone wants to scrape my profile to make a report on industry trends, I’m fine with it. What I don’t want is hiQ vacuuming up my data so they can snitch to my employer if they think I’m job hunting. How is this a paradox? Tech — the web in particular — is supposed to be an equalizing force, but HiQ is clearly trying to give m…

>Tech — the web in particular — is supposed to be an equalizing force

Don't take the Google PR so seriously, the tech industry wants to make money like every other industry.

Re: Congrats! Web scraping is legal! (US precedent)

#152
post #89

Earlier quoted context omitted.

Sure. My data is still my data, and if I publish it on my platform for free, that still shouldn't automatically give you the right to copy the data and provide on your platform. It's basically the same as a TV broadcasting a film for free, and then going after you legally if you recorded that film and uploaded it to your website.

It's not the same thing at all. The film is broadcast without re-transmission rights.

How is that different? Where on my website I gave anyone "re-transmission" or "re-publishing" rights?

Re: Congrats! Web scraping is legal! (US precedent)

#153

>> Most importantly, the appeals court also upheld a lower court ruling that prohibits LinkedIn from interfering with hiQ’s web scraping of its site. This fundamentally changes the balance of power in dealing with such cases in the future. This, I don't agree with. I agree that it's not fraud to send a bot to scrape public info but the site should every right to block a bot or person.

This doesn't sit well with me either. However, LinkedIn are trying to redefine a fundamental principle of the web, i.e. easy access to publicly available information, and they're doing it simply to protect their commercial interests.

It would be terrible to see the web compromised by a spat between two (in my opinion) scummy companies, and I think we could do with some hard push-back against attacks on the web generally.

Re: Congrats! Web scraping is legal! (US precedent)

#154

This only affects the ninth circuit—which includes the tech hubs San Francisco, Seattle, LA, and Portland. It would only apply to the rest of the country if the Supreme Court affirmed it. Even then, a well-funded company or zealous prosecutor could say that it doesn’t apply in your case because of some technicality. In that case you would need hundreds of thousands or millions of dollars and a few years to litigate t…

Is that how really circuit court rulings get applied? I always understood each ruling on the rungs up the ladder to the supreme court applied across the land until a final ruling was determined.

Circuit court rulings are usually only binding precedent within their district. However, the Court of Appeals for the Federal Circuit has exclusive appellate jurisdiction over certain subject matters (e.g., patents), so it's supposed to follow appropriate precedent for stuff outside its remit and its precedent is binding on everybody for stuff inside its remit.

That said, it is not unusual for a court to look to rulings in other jurisdictions to decide a matter if there is no binding precedent in place. They are not required to, however.

Re: Congrats! Web scraping is legal! (US precedent)

#155
post #119

"HiQ only takes information from public LinkedIn profiles. By definition, any member of the public has the right to access this information. Most importantly, the appeals court also upheld a lower court ruling that prohibits LinkedIn from interfering with hiQ’s web scraping of its site." Surely I'm not reading this correctly. This would seem to suggest that websites are not legally allowed to prevent bots from crawli…

> Surely I'm not reading this correctly. This would seem to suggest that websites are not legally allowed to prevent bots from crawling their sites. Lots of sites have ToS preventing such things, are those legally void now? Are captchas on public pages illegal, even if you request the page 8000 times in a second? This is just a preliminary injunction. This wasn't an actual ruling on the case. This just says that unti…

You don’t understand what a preliminary injunction is then.

It’s a very, very strong indication that they will win. Courts don’t issue preliminary injunctions unless it’s extremely likely the side who won the preliminary injunction will win.

Re: Congrats! Web scraping is legal! (US precedent)

#156

Earlier quoted context omitted.

What if the organization is one person in an LLC? Do they get rights? If so then a big company can hire a bunch of little LLCs to act as rights-having proxies for any task that requires them.

I'm going to assume you're asking in good faith and try to address the confusion here. The human does get rights, the organization doesn't. In some cases, believing that humans have rights and believing that organizations have rights might lead one to the same action. In those cases, I'd take the action. I wouldn't want to violate a human's rights out of some vindictive dislike of organizations: that's not the point.…

Let's say that individual humans have the right to keep secrets. Let's also say that they have the right to keep secrets with their associates, and to tell them to who they please. Now, doesn't that make it legal for a group of people to keep secrets about you? What about selling them? I just don't see what doing away with the legal fiction of corporate personage would do about Facebook.

Re: Congrats! Web scraping is legal! (US precedent)

#158
This is phenomenal timing for Clearview AI who in the last week was exposed by the NYTimes for working with law enforcement to identity suspects via their database of web scrapped images of individuals.

https://news.ycombinator.com/item?id=22173899

https://news.ycombinator.com/item?id=22083775

Re: Congrats! Web scraping is legal! (US precedent)

#159

Earlier quoted context omitted.

> Lots of sites have ToS preventing such things, are those legally void now? Are captchas on public pages illegal, even if you request the page 8000 times in a second? ToS are subservient to the law; you can (probably) terminate a service account from a user that breaks your ToS, but if the user does not have a service account (as is the case for HiQ, it doesn't seem they were using accounts for it), then your ToS do…

I'm trying to puzzle out how this works in practice. So if LinkedIn has truly public data (no login required to view) then it can be scraped no problem. But if it's only accessible with a login, then it falls under TOS and they can be blocked?

[deleted]

Re: Congrats! Web scraping is legal! (US precedent)

#160

"HiQ only takes information from public LinkedIn profiles. By definition, any member of the public has the right to access this information. Most importantly, the appeals court also upheld a lower court ruling that prohibits LinkedIn from interfering with hiQ’s web scraping of its site." Surely I'm not reading this correctly. This would seem to suggest that websites are not legally allowed to prevent bots from crawli…

https://en.wikipedia.org/wiki/Tortious_interference This would mostly mean that you cannot start interfering with webscraping you previously allowed merely because you learned that they're making money with the scraped data.

If I decide to change the class names, or HTML structure of the page, is that no longer allowed?

How far does this go?

Post reply on HN