Live data from Hacker News

Congrats! Web scraping is legal! (US precedent)

parsers.me

251–260 of 409 posts

Re: Congrats! Web scraping is legal! (US precedent)

#251

Earlier quoted context omitted.

It seems absurd if the 'interference' only directly affects their own property. Like, if my neighbors start monetizing livestreaming my backyard, suddenly I can't put up a fence? Except worse because in actuality, this third-party contract is costing them money through server load and bandwidth.

Your analogy doesn't hold. Your backyard is private property. The data that LinkedIn publishes is intended for the public. That's why Google can index the pages and give you results from LinkedIn.

> Your analogy doesn't hold.

It does, in the US. You're likely making an inconsistent comparison.

Property ownership has nothing to do with visual access. You cannot legally be barred from casually (involuntarily) perceiving something. It's reasonable to put up physical barriers to reduce what is casually perceived. It's a very good analogy.

Re: Congrats! Web scraping is legal! (US precedent)

#252
post #237
post #213

Earlier quoted context omitted.

The issue here for some, if not many, is a matter of scale. It is one thing if an end-user, whom I am trying to service, comes to my site and gets my publicly available data. Maybe I monetize with ads, maybe not. It doesn't matter, that is the audience I am trying to service, regardless of size. But when you scrape it my load goes up dramatically. A load I have to pay for. It is analogous to the privacy debates going…

Provide an API for public data to reduce the costs associated with rendering a full blown page, and deliver just the information needed.

Who pays for that API and the bandwidth? What’s in it for the data provider? On LinkedIn, viewing the data now shows ads or at least prompts the viewer to join the network. With scrapping and free API access, how exactly does LinkedIn benefit for their work of hosting the data?

Re: Congrats! Web scraping is legal! (US precedent)

#253

Earlier quoted context omitted.

It seems absurd if the 'interference' only directly affects their own property. Like, if my neighbors start monetizing livestreaming my backyard, suddenly I can't put up a fence? Except worse because in actuality, this third-party contract is costing them money through server load and bandwidth.

Your backyard can be a walled garden--this is about the public front of the property.

Exactly your backyard is of course yours. But you are not at liberty to use it to damage others. There's lots of rules about this. For example, opening a brothel on your own land is definitely not legal without considering how it affects the neighborhood.

Re: Congrats! Web scraping is legal! (US precedent)

#254

Earlier quoted context omitted.

I want to be able to use LinkedIn to network with colleagues and people in my industry. If someone wants to scrape my profile to make a report on industry trends, I’m fine with it. What I don’t want is hiQ vacuuming up my data so they can snitch to my employer if they think I’m job hunting. How is this a paradox? Tech — the web in particular — is supposed to be an equalizing force, but HiQ is clearly trying to give m…

We haven't solved that in the same way we haven't "solved" encryption not having a magical good people only door despite spook tantrums. There fundamentally isn't a possible mechanism and really wanting it doesn't change that. It is a result of equality - not of outcome but rules. Open for everyone but those whose applications you don't like isn't open. On a technical level trying to prevent it is like the "evil bit"…

Of course there are possible mechanisms. There are heuristics to detect bots. The whole reason for this lawsuit is that LinkedIn blocked hiQ from scraping their website.

I'm also not necessarily talking about a technical defense against unwanted scraping. Write a law makes it illegal to do something like "scraping personally identifiable information and storing or presenting it non–anonymized", and prosecute companies who break it. I'm sure there are loopholes in that particular example, but the point is we can absolutely add shades of gray here.

> Open for everyone but those whose applications you don't like isn't open.

Openness should be a means, not an end. If we make something "not open" but it prevents 95% of undesirable uses and only 5% of desirable ones, is that not a tradeoff worth discussing?

Re: Congrats! Web scraping is legal! (US precedent)

#255

Earlier quoted context omitted.

It seems absurd if the 'interference' only directly affects their own property. Like, if my neighbors start monetizing livestreaming my backyard, suddenly I can't put up a fence? Except worse because in actuality, this third-party contract is costing them money through server load and bandwidth.

Your analogy doesn't hold. Your backyard is private property. The data that LinkedIn publishes is intended for the public. That's why Google can index the pages and give you results from LinkedIn.

It's trivial to fix that - the exterior of GP's house then. That's available for public viewing; is intended for it, but is private property. If you monetise livestreaming it and describe it in your ToS, GP can't repaint the front door, or get new windows?

Or perhaps slightly less contrived:

If I publish a monthly lowlights reel of my favourite sports team as a podcast discussion on where they can improve in all their lost games, and then they suddenly go on a winning streak for >1month so my USP is gone and I have nothing to talk about..?

Re: Congrats! Web scraping is legal! (US precedent)

#256

"HiQ only takes information from public LinkedIn profiles. By definition, any member of the public has the right to access this information. Most importantly, the appeals court also upheld a lower court ruling that prohibits LinkedIn from interfering with hiQ’s web scraping of its site." Surely I'm not reading this correctly. This would seem to suggest that websites are not legally allowed to prevent bots from crawli…

Here is my attempt to draw an analogy.

There is a large banner next to the highway that shows some weather information that if properly organized (lets say monthly almanac) you would find people to pay money for it. The banner owner does not make money this way - he ask you to go to his website and signup for an account. But you drive the highway (internet) every day, look at the banner, write down the weather updates, and then offer them on your website as a sale. The owner gets angry and sue you. The court decides you are free to drive by the highway and you free to put your eyeballs on their weather banner, especially given the banner is available to everyone (LinkedIn profiles are avail to view without needing an account) and you are free to use the information you obtained for free without interference with said banner in a form of a monthly almanac that you sell. At the end of the day, the banner owner does not own the weather information that someone else put in there (for example a weather meteorologist).

I think personally its a healthy decision. Otherwise it would be similar to prejudice of who should be allowed to enter and browse a street store that by law is available to everyone.

Re: Congrats! Web scraping is legal! (US precedent)

#257

Let’s not pretend this is a pure win. There are good uses of web scraping, like Archive.org trying to preserve the web. But what HiQ is doing is looking at public LinkedIn profiles and then snitching to employers if they think an employee is searching for a new job. It’s easy to blanket say “web scraping is legal, do what you will“. The tricky part is protecting people’s public data while not giving a huge moat to gi…

That's the thing. Web scraping isn't really the problem here. It's what companies are doing with personal information. If LinkedIn started doing the same thing as HiQ, it would be just as bad (probably worse), but the legality of web scraping is irrelevant to that.

[deleted]

Re: Congrats! Web scraping is legal! (US precedent)

#258

"HiQ only takes information from public LinkedIn profiles. By definition, any member of the public has the right to access this information. Most importantly, the appeals court also upheld a lower court ruling that prohibits LinkedIn from interfering with hiQ’s web scraping of its site." Surely I'm not reading this correctly. This would seem to suggest that websites are not legally allowed to prevent bots from crawli…

> hiQ argued

That does not mean that hte court agreed.

The judges said that CFAA doesn't apply.

In other words, the judges said that LinkedIn couldn't use the US legal system to force HiQ to stop. Judges didn't say that LinkedIn was barred from using technical measures.

The court did allow a preliminary injunction against LinkedIn, due to the possibility of "monopolies" (to be determined in Court later), pending resolution of that latter question.

LinkedIn might still win their claim to their right to block scrapers via technical means.

Re: Congrats! Web scraping is legal! (US precedent)

#259
post #119

Earlier quoted context omitted.

> Surely I'm not reading this correctly. This would seem to suggest that websites are not legally allowed to prevent bots from crawling their sites. Lots of sites have ToS preventing such things, are those legally void now? Are captchas on public pages illegal, even if you request the page 8000 times in a second? This is just a preliminary injunction. This wasn't an actual ruling on the case. This just says that unti…

You don’t understand what a preliminary injunction is then. It’s a very, very strong indication that they will win. Courts don’t issue preliminary injunctions unless it’s extremely likely the side who won the preliminary injunction will win.

Please read the linked decision before putting words in judges' mouths:

https://parsers.me/appeal-from-the-united-states-district-co...

Re: Congrats! Web scraping is legal! (US precedent)

#260
post #229

"HiQ only takes information from public LinkedIn profiles. By definition, any member of the public has the right to access this information. Most importantly, the appeals court also upheld a lower court ruling that prohibits LinkedIn from interfering with hiQ’s web scraping of its site." Surely I'm not reading this correctly. This would seem to suggest that websites are not legally allowed to prevent bots from crawli…

So if I (like many others) have cloudflare web scraping protection turned on, that is now against American law?

There's no reason to believe that.
Post reply on HN