Live data from Hacker News

Congrats! Web scraping is legal! (US precedent)

parsers.me

201–210 of 409 posts

Re: Congrats! Web scraping is legal! (US precedent)

#201
So, the whole Google's search is based on scraping publicly accessible information from public web-sites. Nothing to be surprised about that the practice is declared legal. Otherwise they should have demand for Google to basically liquidate Google search unless it signs an agreement with every single site in the index.

Re: Congrats! Web scraping is legal! (US precedent)

#202
post #143

Earlier quoted context omitted.

In this case, LinkedIn users kind of do want their “public profiles” to be public. They’re online CVs; by definition, if you make one, your goal is to get it into the hands of anyone who asks for it! LinkedIn, likewise, has built its business model on an implicit contract with its users that it’s going to show their CV to anyone who asks for it. I think LinkedIn users would be surprised that LinkedIn doesn’t let bots…

>I think LinkedIn users would be surprised that LinkedIn doesn’t let bots read (scrape) their public CV. LinkedIn is really really clear that: - They won't share you information with 3rd parties - You're not allowed to use information on LinkedIn for commercial purposes without their permission - Other users can view your personal data So, why would I expect random third party companies to be able to scrape and sell…

There’s a difference between “my information” and “the public webpage that I went through a publishing workflow to create from a curated selection of my information.”

Let me put it this way: if I have a Wordpress blog, I’d certainly be miffed if Wordpress let bots see my drafts... but I’d also be miffed if Wordpress didn’t let bots (Google, for one!) see the published blog itself. It’s a blog; a public website! Anyone or anything with the URL is supposed to be able to retrieve the page! It’s not “my information” any more†; it’s been broadcast!

† You might want to mentally analogize to copyright, but I don’t think it’s the right model for the intuition people have here. Instead, try mentally analogizing to confidentiality. When a classified document is published in the public sphere (e.g. as evidence in a trial, as testimony before congress, etc.), this forcibly declassifies it. No matter how much the originator of the document might want to still keep it a secret, the legal protections of confidentiality don’t apply to it any more: it’s out there now. Anyone who reads it could plausibly have just read the public-sphere copy, so there’s no longer any way to charge people who have knowledge of the previously-classified information with any crime.

Re: Congrats! Web scraping is legal! (US precedent)

#203

Earlier quoted context omitted.

There's a long, long history (probably hundreds –if not thousands– of years old) of selling aggregated or processed publicly-available information. I'm not particularly thrilled with it, but enough people think of it as a valuable enough service to pay for; even if they know they could get it themselves, for free. LinkedIn users (as opposed to the company) might actually like what HiQ is doing, as it may help their o…

> but enough people think of it as a valuable enough service to pay for; even if they know they could get it themselves, for free. It's not free, it takes time to collect data. Buying it makes a lot of sense as long as you pay less than what's your own time worth to you...

It is true in the current situation, though I would prefer that we ensure free data must be free. In that case buyers of data would be incentivized to pressure providers of free data to improve the data quality.

Re: Congrats! Web scraping is legal! (US precedent)

#204

Earlier quoted context omitted.

https://en.wikipedia.org/wiki/Tortious_interference This would mostly mean that you cannot start interfering with webscraping you previously allowed merely because you learned that they're making money with the scraped data.

It seems absurd if the 'interference' only directly affects their own property. Like, if my neighbors start monetizing livestreaming my backyard, suddenly I can't put up a fence? Except worse because in actuality, this third-party contract is costing them money through server load and bandwidth.

Your analogy doesn't hold. Your backyard is private property. The data that LinkedIn publishes is intended for the public. That's why Google can index the pages and give you results from LinkedIn.

Re: Congrats! Web scraping is legal! (US precedent)

#205

Earlier quoted context omitted.

One of my clients is involved in property tax collection and reporting. Property Tax records are public info, and their website allows looking up the records for any property without a login. However, the data behind this website it the _source_ of the public records, and not the public records themselves (which would be local government databases). For years now we've been in an arms race with someone using a botnet…

Maybe you could contact the scrapper? Just post magnet links on the site that allows them to get nicely formatted dump of what they want.

We did figure out who the scraper probably is, but only after several years. For a long time they used an untracable botnet, but after blocking that they eventually switched to a corporate network we traced to a data aggregation company. But we don't know for sure who's doing the scraping; it could be the company, a rogue employee, or a botnet that got loose on their network.

Re: Congrats! Web scraping is legal! (US precedent)

#206
post #202

Earlier quoted context omitted.

>I think LinkedIn users would be surprised that LinkedIn doesn’t let bots read (scrape) their public CV. LinkedIn is really really clear that: - They won't share you information with 3rd parties - You're not allowed to use information on LinkedIn for commercial purposes without their permission - Other users can view your personal data So, why would I expect random third party companies to be able to scrape and sell…

There’s a difference between “my information” and “the public webpage that I went through a publishing workflow to create from a curated selection of my information.” Let me put it this way: if I have a Wordpress blog, I’d certainly be miffed if Wordpress let bots see my drafts... but I’d also be miffed if Wordpress didn’t let bots (Google, for one!) see the published blog itself. It’s a blog; a public website! Anyon…

Well, my name, my job title, my employer, my job history.

These are all my information, and selling them to marketing companies is definitely not archiving.

Would you be OK with a company scraping your blog and selling it?

Re: Congrats! Web scraping is legal! (US precedent)

#207
post #119

Earlier quoted context omitted.

> Surely I'm not reading this correctly. This would seem to suggest that websites are not legally allowed to prevent bots from crawling their sites. Lots of sites have ToS preventing such things, are those legally void now? Are captchas on public pages illegal, even if you request the page 8000 times in a second? This is just a preliminary injunction. This wasn't an actual ruling on the case. This just says that unti…

You don’t understand what a preliminary injunction is then. It’s a very, very strong indication that they will win. Courts don’t issue preliminary injunctions unless it’s extremely likely the side who won the preliminary injunction will win.

The issuance of an injunction is in no way related to how the future court battle will result.

Re: Congrats! Web scraping is legal! (US precedent)

#208
post #22

Earlier quoted context omitted.

See that's where I have problem with this. Isn't data just _data_? Lets draw some pararells to real life. If I go to public space like town square - can't I take pictures, notes and records then go home and draw my analytics from it? What if I read something in a book I bought, can't I quote it? Same thing should be with web resources even if they are creative - as long as I don't publish them I should be able to scr…

This is why I strongly prefer the Dutch term 'auteursrecht' (author's rights) as opposed to copyright. Copyright has this annoying incorrect connotation that it has anything to do with copying when it's really publishing that it should be limiting. Downloading publicly available data should (by definition of public) not be a violation of someone's rights. However it's easy to see why it wouldn't be desirable for some…

>Copyright has this annoying incorrect connotation that it has anything to do with copying when it's really publishing that it should be limiting. //

Copyright does make _copying_ tortuous. Broad personal use exceptions in USA, for example, make this appear not to be true, but it is the act of copying - even without publication - that is protected in general.

Ripping a CD in UK, for example is copyright infringement without a general personal use exception (there are exceptions, under Fair Dealing, but whatever you're doing almost certainly doesn't fall into them).

See eg UK CDPA1988, Chapter II, section 16(1)(a); or USC17, Chapter 1, 106(1).

Re: Congrats! Web scraping is legal! (US precedent)

#209

This only affects the ninth circuit—which includes the tech hubs San Francisco, Seattle, LA, and Portland. It would only apply to the rest of the country if the Supreme Court affirmed it. Even then, a well-funded company or zealous prosecutor could say that it doesn’t apply in your case because of some technicality. In that case you would need hundreds of thousands or millions of dollars and a few years to litigate t…

Is that how really circuit court rulings get applied? I always understood each ruling on the rungs up the ladder to the supreme court applied across the land until a final ruling was determined.

Across the land within their circuit, over matters within their jurisdiction. Elsewhere the ruling is merely advisory in nature.

It gets tricky with nationwide actors though.

Besides some specialized topics like patents and international trade, nationwide orders and injunctions are the sort-of exception, which are based on a courts local jurisdictional power over a non-local nationwide actor.

I'm not sure what the generalized name of the principle (beyond injunctions) is called (I've heard it as a type of jurisdictional overreach), but the presumption as applied here is that LinkedIn, being in the 9th circuit's jurisdiction, also adhere to the ruling outside the circuit, absent a contradictory ruling by a different circuit.

(One of the usual requirements for the Supreme Court to even hear a case is that different circuits have conflicting rulings on a matter.)

But it looks like the Supreme Court is about to seriously reign that nationwide power back soon, at least for judges issuing orders to departments/actors of the executive branch.

Re: Congrats! Web scraping is legal! (US precedent)

#210
post #128

Earlier quoted context omitted.

One of my clients is involved in property tax collection and reporting. Property Tax records are public info, and their website allows looking up the records for any property without a login. However, the data behind this website it the _source_ of the public records, and not the public records themselves (which would be local government databases). For years now we've been in an arms race with someone using a botnet…

Would rate limiting be a viable solution?

We've done that, but it's tough to rate-limit a botnet because of the ip address spread. Also, their crappy scraper software doesn't even bother to check if requests are successful; it spews them just as fast no matter how our site responds.
Post reply on HN