Live data from Hacker News

LinkedIn: It’s illegal to scrape our website without permission

arstechnica.com

271–280 of 303 posts

Re: LinkedIn: It’s illegal to scrape our website without permission

#271
post #211

It is illegal to let your browser fetch this text without my permission.

I suspect we don't need your permission as under the terms of use of this sight you forfeit your rights to this user generated content, and the rights now belong to Hacker News.

Re: LinkedIn: It’s illegal to scrape our website without permission

#272

As far as I'm concerned, if the information is publicly visible on a site (not behind a login) and as long as the scraping doesn't cause performance issues or generate costs on the site, then it should be perfectly fine to do.

Whilst that is your opinion, it is not a fact, nor is it legally valid.

Re: LinkedIn: It’s illegal to scrape our website without permission

#273
post #100

Earlier quoted context omitted.

Actually, no. It's like taking pictures of the paintings from the street, and reselling those pictures. If they don't like that... that the painting down. Simple.

Hmm, then that becomes a copyright issue I suppose.

Well its a Database Right, which is a property right rather than copyright.

https://en.wikipedia.org/wiki/Sui_generis_database_right

As with all rights, it varies with jurisdiction.

Re: LinkedIn: It’s illegal to scrape our website without permission

#274
LinkedIns Email Password ConArtistery has inspired a whole generation of Anti-Malware and Lawyer-Plugins, preventing the layman from giving away his data, even on "friendly" sites.

Im going to turn around now, and whatever happens to this site is going to happen. They worked so hard, to ask for this.

Re: LinkedIn: It’s illegal to scrape our website without permission

#275
> To expand its user base, Power asked users to provide their Facebook credentials and then—with their permission—sent Power.com invitations to their Facebook friends. Facebook, naturally, didn't appreciate this marketing tactic. They sent Power a cease-and-desist letter and also blocked the IP addresses Power was using to communicate with Facebook's servers.

> Facebook sued, claiming that its cease-and-desist letter made Power's access unauthorized under the terms of the CFAA. Power disagreed and argued that having permission from Facebook users was good enough—it didn't need separate approval from Facebook itself.

How can be illegal if users are giving their permission? What happens if I give my permission to an external service to extract my own data?

Re: LinkedIn: It’s illegal to scrape our website without permission

#276
post #251
post #230

Earlier quoted context omitted.

That argument works, insofar as it does, only for more recognizable bots and browsers. If I write a client of some sort that identifies itself as: Mozilla/5.0 (Windows NT 6.1) AppleWebKit/537.36 (KHTML, like Gecko) Snackmaster Pro/666.0.666 What do you do? I also tell my browser to lie about what it is sometimes, due to sites that are malfunctioning, but whose owners choose to document the errors instead of fixing th…

In your first case, if you are running on Windows NT 6.1 using WebKit on a new browser for humans called 'Snakemaster Pro', then you aren't doing anything wrong. If by client you mean a robot, then you are pretending to be a browser and you are accessing the service without permission. Let me ask you a question, say your client was hitting my service with that user agent, 100 times a second, crawling through urls seq…

Have you ever heard of "headless browsers" (like [chrome](https://github.com/dhamaniasad/HeadlessBrowsers/issues/37)? What are some defining characteristics of browsers that are absent in scraping clients? If I open a browser window while doing the scraping is that acceptable?

Re: LinkedIn: It’s illegal to scrape our website without permission

#277

Earlier quoted context omitted.

I disagree with your analogy. To me, the key word in "HTTP request" is request . A request is something that can be granted or not.

Perhaps more clearly would be any HTTP request that LinkedIn believes is in violation of their terms of service will be denied. It can be hard to know when the first request arrives if it is someone scraping the site or not, but once it is clear that it is someone scraping they actively deny all future requests. If they could know that the request coming in was going to be a scrape and not a page view they would pree…

But what is the difference between a scrape and a page view? If a human looks at it once, after scraping, does it become a page view? Is pocket downloading content on my behalf for me to read later, a scraper? What's the difference between a scraper and an offline browser who's content a human never browses?

Re: LinkedIn: It’s illegal to scrape our website without permission

#278

Earlier quoted context omitted.

> Am I wrong? I hate to be the bearer of bad new but...maybe. A reading of the Computer Fraud and Abuse Act could make robots.txt legally enforceable. And given the government's approach to CFAA cases a very aggressive interpretation, under the right circumstances (for example, when it provides evidence that the scraper knew that scrapint was not authorized), seems like a real possibility. Among the many other things…

I wouldn't be too fast to jump to conclusion robots.txt legally enforceable. You would need to cite prior case law's. Without any case law's it make decision on a law error prone at best.

The possibility still exists. After reading this (insanely broad) definition I think the chance is not even that low.

Re: LinkedIn: It’s illegal to scrape our website without permission

#279
post #88

Earlier quoted context omitted.

The number of photos is irrelevant to the analogy, though, as is what people do with the photos afterwards. If the bikes are visible from the public street, people can take as many pictures of every bike they want, and then make money from them if they want. It doesn't affect the owners' usage of the bike (unlike the original analogy, where the owner loses access, which was what I was trying to correct) Physical anal…

> The number of photos is irrelevant to the analogy Actually, the size of the data and the number of requests is very relevant. More data means more information, means more money. It also means more bandwidth and processing power required to process requests. You're not taking a photo of the bike, you're asking the bike to give you a photo of it. > it's just dishonest/misleading to pretend that copying data is ever a…

If LinkedIn are being that negatively affected by a single scraper, they should deal with it - block it, only allow a specific number of requests from an IP per day, anything that doesn't involve lawsuits. The problem is them trying to pretend that publicly visible content is really private if they say so, without them trying to protect it in any real way.

"The hurt occurs when people benefit from the work the original author put into creating that data without proper compensation"

Not necessarily. If I'm paying for print of some imaginative artwork that was created using the picture of the bike, that doesn't mean the bike owner lost anything, even if he spent time building the bike with his own hands. Similarly, if the only reason why people paid Hi-Q was for the extra work that they put in, LinkedIn didn't lose money because people would not have bought their product without that extra work.

There is certainly an argument that Hi-Q should have licenced the content first, but it's public data. If they want to make licence deals, don't put it in the view of the public street then whine when people are documenting what's in public.

"It's a straw man."

No, the straw man is pretending that a copy is the same as theft. Theft is theft because someone is depriving you of the original, not because you imagine you might have had more sales if the copy didn't exist. There's a reason why there are different words for different things, and pretending that a copy is the same as taking a physical object it a lie. Period.

"I place hours of working into something that doesn't put food on the table because you can clone my work, but I can't clone my food."

But, you put the price up too high, so I opted not to buy it. Maybe borrow the CD from a friend, or listen to something else. Or, you decided I couldn't buy it in the format or region I wanted. There are real issues, but pretending that a copy = a lost sale is utter bull that's been debunked time and time again, yet is regularly repeated by people trying to inject emotional arguments instead of facts.

"I'd say it's pretty obviously interfering with their business model"

Then perhaps they should address the business model or not put their content out there in public unprotected if it's that valuable to their income.

"LinkedIn could ban IPs that make unreasonable number of requests in a short amount of time."

Yes they could. Which would not have to involve the courts in any way. Or, they could protect the content in some other way that (for example) requires a log in and adherence to T&Cs, with which they could easily kick violators off their site for non-compliance.

The issue is that LinkedIn are trying to have it both ways - gathering the benefits of public content while blocking others who use the now-public content in ways that are usually acceptable for public content to be used. Sorry, not acceptable, you pick one - take the content away from the public street or accept that some people will use what has been shown to the public.

Re: LinkedIn: It’s illegal to scrape our website without permission

#280
post #28

If you post it on a public network, it is defacto, public. If you don't want it scraped, take it down, or put it behind a login. If the user provides the login to a scraper, then the scraper has permission.

That's a pretty literalist, one-size fits all approach to policy. I don't think it's a good framework to use for applying ethics considerations. If I can walk near a pool, should I also be able to run? Is running anything more than faster walking? If I'm allowed to be around the pool walking with my entry ID, should I also be allowed to place my ID on a little motorized car and make it dart around the pool really fas…

What LinkedIn is doing seems more like a pool threatening to sue people that look at the facility itself, from outside of their property. It's not clear whether viewing a public website is more (reasonably) similar to entering someone's property versus looking at their property from outside of it. To me, it seems much more like looking at it from outside. But, fortunately for LinkedIn, they can prevent people from viewing their site! (They just have to figure out for whom they want to do so.)
Post reply on HN