Live data from Hacker News

LinkedIn: It’s illegal to scrape our website without permission

arstechnica.com

1–10 of 303 posts

Re: LinkedIn: It’s illegal to scrape our website without permission

#4
A few weeks ago I was looking up public profiles on LinkedIn and I noticed what I interpreted to be some network-side fingerprinting of some kind. The first couple profiles came up, but from that point I was only served a sign up page. It didn't matter if I changed browsers or spawned new incognito sessions.

Re: LinkedIn: It’s illegal to scrape our website without permission

#5
>One plausible reading of the law—the one LinkedIn is advocating—is that once a website operator asks you to stop accessing its site, you commit a crime if you don't comply.

I'm fine with that.

But I'm also fine if they do the like the article suggests and require anything non-scrapable to be behind an account prompt, even if everyone with a account can access it.

I don't think it's fair to make Linked In foot the bill for someone else's business. They shouldn't have to serve that content to people who aren't actually their users.

Re: LinkedIn: It’s illegal to scrape our website without permission

#8
post #6

"illegal"? Under what law?

2nd paragraph, FTA:

"this scraping violated the Computer Fraud and Abuse Act, the controversial 1986 law that makes computer hacking a crime. HiQ sued, asking courts to rule that its activities did not, in fact, violate the CFAA."

Re: LinkedIn: It’s illegal to scrape our website without permission

#9
post #2

Not sure how to feel about this. While in theory, scraping data is a shady practice - companies like LinkedIn leave the door open for it.

Ironic because LinkedIn scrapes and collects every scrap of contact info they can find. I've got people I sent an email to once in 199? suddenly popping up suggesting we know each other.

Re: LinkedIn: It’s illegal to scrape our website without permission

#10
I'm following this story with a lot of interest. I've done (and still do!) a lot of data crawling/scraping. In the past I've worked on so-called "alternative data" collection and analysis for financial forecasting.

Without going into too much detail, a lot of hedge funds have teams constantly searching for kernels of data that can contribute some kind of signal for market movements. This data can come in the form of satellite imagery for oil tankers or manufacturing centers, but it can also come from the very creative use of scraped and aggregated data. It's typically very difficult to identify, collect and analyze on a technical level (as 'chollida1 has lamented in the past: normalization, labeling/bucketing and analysis of disparate data across different formats, sources and processing timeframes is a pernicious problem at this scale). From a compliance standpoint there are also generally strict requirements governing legality of use.

Depending on the specific data, you might be capable of predicting earnings or broader market movements with a hiQ Labs doesn't collect data for this specific purpose, but it is absolutely related. In the past I have stayed away from crawling LinkedIn and Yelp precisely because they are very litigious (regardless of the eventual outcome and legality). Now that there's another relatively high profile case out in the open like this, I'm interested in seeing how it proceeds and what the ramifications will be for companies that collect data across a wide range of uses. As Grimmelman mentioned in the article, this can impact a lot of types of businesses, not just those in the same space as hiQ. Outside of finance I am familiar with many tech companies which (openly or otherwise), kickstarted what are now widely known enterprises through cleverly crawling or scraping massive amounts of data.

Post reply on HN