Earlier quoted context omitted.
What if the user is provided with permission to view medical records - your medical records for example? Would it be ok to scrape and sell those? What about children's school profiles and records?
See the first line. If there is a problem with that, see the second line
LinkedIn: It’s illegal to scrape our website without permission
131–140 of 303 posts
Re: LinkedIn: It’s illegal to scrape our website without permission
#132Earlier quoted context omitted.
People have gone to jail for it. https://www.wired.com/2010/11/wiseguys-plead-guilty/
I'd say "buy up all tickets with fake data for scalping" weighted in more than just passive scraping.
The EFF said this, so they seem to agree that it set a bad precedent, even for those just scraping.
"Under the government's theory, anyone who disregards – or doesn't read – the terms of service on any website could face computer crime charges," said EFF civil liberties director Jennifer Granick in a press release at the time. "Price-comparison services, social network aggregators and users who skim a few years off their ages could all be criminals if the government prevails."
Re: LinkedIn: It’s illegal to scrape our website without permission
#133I'm following this story with a lot of interest. I've done (and still do!) a lot of data crawling/scraping. In the past I've worked on so-called "alternative data" collection and analysis for financial forecasting. Without going into too much detail, a lot of hedge funds have teams constantly searching for kernels of data that can contribute some kind of signal for market movements. This data can come in the form of…
Well, it's pretty simple to legally dodge these kinds of threats. If you scrape regularly, then pick up a dozen or more machines around the world, in less than friendly areas to US law. Pay with a rechargeable credit card or bitcoin. And the buy servers and set up a hadoop cluster that handles scan-jobs. The worst case scenario is that LinkedIN, YELP, and others get some of your servers shut down. Wash, rinse, repeat…
Re: LinkedIn: It’s illegal to scrape our website without permission
#134Earlier quoted context omitted.
LinkedIn has two main functions: 1., Self-updating rolodex for sales people. 2., Recruitment tool. It works for both, no real competitor in sight due to the massive network effect. Some niche networks are doing ok, like Xing in DACH region. Not aware of anything special in China, guess everyone is on WeChat anyhow.
This is not what they are offering, look at the premium offerings: https://premium.linkedin.com/ do you know the ROI of using LinkedIn Premium (e.g. InMail) to just contact people using a zillion of methods available, and then adding them to your LinkedIn...
below the marketing language this is exactly those 2 points. guess it is easier if you work in enterprise, this business speak is a different language. pretty verbose, low information density.
Re: LinkedIn: It’s illegal to scrape our website without permission
#135Earlier quoted context omitted.
This is not what they are offering, look at the premium offerings: https://premium.linkedin.com/ do you know the ROI of using LinkedIn Premium (e.g. InMail) to just contact people using a zillion of methods available, and then adding them to your LinkedIn...
not sure where the confusion is. below the marketing language this is exactly those 2 points. guess it is easier if you work in enterprise, this business speak is a different language. pretty verbose, low information density.
Re: LinkedIn: It’s illegal to scrape our website without permission
#136Earlier quoted context omitted.
This sounds like the setup to an up-spiraling arms race with "dark scrapers". I'm guessing behind-the-scenes, LinkedIn figures they can out-spend the extralegal scrapers, and likely considers their efforts will deliver halo effects to the rest of Microsoft. It would be educational to hear how LinkedIn plans to take down bot-net-based scraping that uses deep learning to identify patterns that successfully mimic human…
I'm skeptical of the potential of dark scrapers at scale. You'd need to simulate too much human behavior to be unidentifiable, and humans are slow. You would need real-looking bot accounts that you'll use to scrape. You'd need a realistically randomized rate limit, sampling from some distribution conditional on the type of the source page. You'd need realistic mouse/keyboard movements. Realistic hours of operations.…
Average botnet size is 20,000 compromised PC's. Srizbi is estimated at 450,000. Another vector I'd explore is teaming up with crypto-miners. As I understand it, there are no economic returns tapping into the CPUs any longer, so miners are using only GPUs and ASICs; if this is true, they'll have some spare CPU cycles, that they'd probably be willing to rent out to get some marginal returns on the CPUs that have to run and manage the mining chips, running a JVM or some other VM. If we can do that, then we can probably tap 2-3M hosts, many of them rotating in and out per day.
Throw out an army of mechanical turk assignments to get real humans to register fake accounts. They get paid upon submitting an account and password, which your scraping servers verify, then change the password and commandeer. Perhaps have them register the fake account while running under a container or VM on their computer; the container/VM is instrumented to capture all activity. The activity metrics and data are uploaded to a deep learning system, that identifies the patterns that work and the ones that don't, and uses that to guide the developers of what to randomize, and by how much.
Add in a component to randomly invite/follow other fake and real accounts, and generate Markov-chain-generated copypasta. Set aside a portion of the fake accounts to only build up networks of users. Initially restrict the market of customers to those who only want once-a-year-updated data. As the network builds, use the notification of changes to selectively scrape only changed user profiles, and upsell for more up-to-date profiles at that time.
If I was LinkedIn, I'd probably concentrate on infiltrating botnet operators, and shutting them down. It would be one large cat-and-mouse game.
Re: LinkedIn: It’s illegal to scrape our website without permission
#137The last one, I did 13 crawlers to keeping comparing prices of all drugstore products online. I'm selling it to drugstore ecommerces who wants to know when their competitor prices and when they're doing promotions, if competitors prices are higher... Well, I just automated the job of a person that was doing this job manually every single day, looking into competitor's site
If it is illegal, you have to ask to google maps to remove all houses from your database and just let in the houses who gave the permission to it.
The both sentence means the same bullshit, they are both wrong:
Google Maps: "The front-door of my house is faced to a public street. That does not mean you can take photos of it on Street View and use a very smart OCR, to read my house number. Not to mention that sometimes you give my house photo to others by Captcha asking the house number, c'mon!"
Linked: "My CV is half-public. That does not mean you can take crawler of it."
... This is only a cool discussion in North Korea or maybe in China. Not in the rest of the world.
Re: LinkedIn: It’s illegal to scrape our website without permission
#138Re: LinkedIn: It’s illegal to scrape our website without permission
#139Re: LinkedIn: It’s illegal to scrape our website without permission
#140Ah, complaining when people look at the painting in your store front gallery on 5th avenue because you have to pay for the upkeep of the sidewalk.
Actually, no. It's more like complaining about others selling tickets to view said painting from the sidewalk. HiQ repackages and sells data it scrapes from LinkedIn.
edit: didn't see the sibling reply which makes the same point