Live data from Hacker News

LinkedIn: It’s illegal to scrape our website without permission

arstechnica.com

261–270 of 303 posts

Re: LinkedIn: It’s illegal to scrape our website without permission

#261

Unpopular opinion: when you make a HTTP request you're asking the server to give you information. The server has the right to say no. IMHO, LinkedIn doesn't have a right to stop scraping after the fact, but they have the right to take technical steps to stop scrapers from accessing their site.

Generally I agree, but with linkedin anything of value seems to require logging in, which means their terms of service come into play. This is very different to scrapping data that they make open and browsable by all.

Re: LinkedIn: It’s illegal to scrape our website without permission

#262

Earlier quoted context omitted.

The pool has the right to kick you out, same as any website. The pool cannot call the police and charge you with a felony for misusing their resources.

The website can't really kick you out though, it can only kick your agent out and you can trivially create a thousand more. The website can politely ask you to stop just like the pool, but it can't actually do anything if you ignore it.

Except block your IP address, or your user agent, or the pattern your software makes when it connects.

Yes, that will cause potential issues for other people, which is why they tend not to do that, but if you trivially create a thousand more agents, and potentially trigger a degradation of service, how are you different to the people who block junctions at traffic lights?

I'm not keen on inconveniencing people, and "it's not that bad" is a poor argument for doing something that someone has explicitly asked you not to do.

Re: LinkedIn: It’s illegal to scrape our website without permission

#263
post #60

Earlier quoted context omitted.

Does the pool have any recourse if you proceed to bypass the ban? Do you have to re-enter to pool to bypass it, or does sending in confederates with their own badges to continue your work also bypassing the ban? How about sending in new motorized cars? The analogy is starting to break down, but I think it's still instructive for the problem of applying a simple first principles approach.

There is a legal concept known as "attractive nuisance"[1]. If I have a pool and neighborhood kids come to play and someone gets hurt, it's my fault. Even if I was away from my house and never gave permission (or explicitly forbade them from swimming), if I don't have proper access controls in place, the courts say it is too tempting for the neighbors to just come over and swim. I need to put up a locking gate to kee…

"An unlocked car is too tempting for some people to just walk past and not take it."

I am saddened by this.

Re: LinkedIn: It’s illegal to scrape our website without permission

#265
post #216

Earlier quoted context omitted.

Alternate data doesn't even need to be as sexy as satellite photos, hell you almost certainly want the data that isn't sexy, the stuff people haven't thought of because it's too boring. Alternative data vendors above all want sales, and even the funds themselves want things to show off to clients. This gives you great opportunities to look at the alternative data they aren't touching. Given this is a predominantly a…

The reason hedgefunds look at satellite images and oil tankers is because everyone looks at sequential ids and price changes so that doesn't give an edge.

That's simply not true - equity analysts can cover anywhere between 5 and 500 stocks, do you really think they have the time or skill set to track all of that? It really is laborious, grueling work.

If you look at the possible returns the equity market is going to make from a stock in a dollar value, and how much research spend is as a percentage of that, you'll quickly see it doesn't pay for much.

You can tell simply by looking at the broker research - that's probably the extent that analysts take things.

The big stocks obviously have a lot of it happening. (eBay listings, airline pricing etc is obviously touted a lot)

But once you start to go down to the mid caps, you enter a void where there isn't much heavy data focused research done, and it's very possible you can have a better gauge of the business than any other investor on the planet once you pull out this data out.

Re: LinkedIn: It’s illegal to scrape our website without permission

#266
post #28

If you post it on a public network, it is defacto, public. If you don't want it scraped, take it down, or put it behind a login. If the user provides the login to a scraper, then the scraper has permission.

That's a pretty literalist, one-size fits all approach to policy. I don't think it's a good framework to use for applying ethics considerations. If I can walk near a pool, should I also be able to run? Is running anything more than faster walking? If I'm allowed to be around the pool walking with my entry ID, should I also be allowed to place my ID on a little motorized car and make it dart around the pool really fas…

The publisher of the information is ultimately responsible for the disclosure of the information. How it is read is of no consequence, as the information has been provided to be read. Certainly, there are issues about resource usage i.e. heavy readers, such as scrapers, but throttling is a perfectly acceptable approach to overuse of resources for both parties, as access is still available over managed resources.

The nature of the original complaint is authorised access to publicly published information. Again, if you do not want people to read publicly published information, do not publish it publicly.

And as a side note - we don't need inappropriate analogies; the web is real, we can discuss the real issue.

Re: LinkedIn: It’s illegal to scrape our website without permission

#268

Unpopular opinion: when you make a HTTP request you're asking the server to give you information. The server has the right to say no. IMHO, LinkedIn doesn't have a right to stop scraping after the fact, but they have the right to take technical steps to stop scrapers from accessing their site.

I don't think you've characterized this accurately. When you make an HTTP request to LinkedIn you are accessing their service. There is a long history of this relationship, you plug your house into the sewer line and you connect to the sewer service. You connect to the power pole and connect to the electricity service. You connect to the telephone pole and connect to the telephone service. Every service has "terms of…

The analogy of pouring waste is not accurate, we are looking at what is done with the service, e.g. it would be the equivalent of allowing drinking the water but not cooking with it, or using electricity for specific devices. The contract is a debit/volume, what I do with it is irrelevant, and by this analogy on the web I should be allowed to scrape if I stay in the allowed bandwidth by the website.

Re: LinkedIn: It’s illegal to scrape our website without permission

#269
post #12

Earlier quoted context omitted.

The CFAA, which makes it a crime to access a computer system without proper permission.

That's a load of controversial hogwash. If you have permission to view the page, then you have permission to scrape the page. Anyone can make a LinkedIn account, ergo anyone can scrape any public profile.

Being able to do so doesn't make it legal.

If you have permission for one thing it doesn't mean you have permission for something different.

Post reply on HN