Earlier quoted context omitted.
I mean it should. That is also a huge anti-competitive action that just isn't pursued by anyone yet: google makes money of scraping and denies scrapers scraping them - that's all sorts of messed up. The problem is that someone would have to sue google first and no one will do that unless there's big business incentive and big business can already scrape the shit out of google. This is the weird thing about web-scrapi…
Google only scrapes sites that allow it by their robots.txt file so I don’t think their policy is as hypocritical as you are making it sound.
9th Circuit holds that scraping a public website does not violate the CFAA [pdf]
181–190 of 293 posts
Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]
#182Earlier quoted context omitted.
I mean it should. That is also a huge anti-competitive action that just isn't pursued by anyone yet: google makes money of scraping and denies scrapers scraping them - that's all sorts of messed up. The problem is that someone would have to sue google first and no one will do that unless there's big business incentive and big business can already scrape the shit out of google. This is the weird thing about web-scrapi…
Google only scrapes sites that allow it by their robots.txt file so I don’t think their policy is as hypocritical as you are making it sound.
Robots.txt isn't something that bars access to information. It's just a notice that the administrator does not want large amounts of queries against certain resources.
Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]
#183Earlier quoted context omitted.
> However, how is it reasonable to force a web site to serve its contents to a third-party company, without being allowed to make a decision whether to serve it or not? Your statement makes absolutely no sense. That's not how internet works. If you serve something publicly you don't get to cherry pick who sees it. Not only it makes no sense technically it's also a huge anti-competitive case.
It makes sense and it is how the internet works. Servers cherry pick who sees their content all the time. Scrapers are often blocked, as are entire IP address ranges. Things like Selenium server scrapers can be (approximately) detected and often are denied access. I’m not sure about being anti-competitive. Serving a website is an action in which you open up your resources for others to access. My friend runs an open…
Could also delay results, offer reduced temporal precision and other things to differentiate use cases.
Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]
#184This action does more than that. The court left the preliminary injunction against LinkedIn in place: "The district court granted hiQ’s motion. It ordered LinkedIn to withdraw its cease-and-desist letter, to remove any existing technical barriers to hiQ’s access to public profiles, and to refrain from putting in place any legal or technical measures with the effect of blocking hiQ’s access to public profiles." So Lin…
Not allowing the CFAA to be (ab)used to attempt to make scraping illegal makes sense. However, how is it reasonable to force a web site to serve its contents to a third-party company, without being allowed to make a decision whether to serve it or not? Serving the web site costs money, and the scraper surely isn't going to generate ad income...
It's actually pretty insane to force a site to serve content. I think both parties are in the wrong here - HiQ for assuming they're entitled to receive a response from LinkedIn's webservers, and LinkedIn for abusing the CFAA to try to deny service rather than figure out a technical solution to their business problem.
In my view:
* The data is public, and free of copyright. If you're a scraper and can get it, you haven't done anything wrong.
* The servers serving the data are still under LinkedIn's control, and they have no obligation or public duty to always serve that content. They could just as well block you based on your IP or other characteristics. If they want to discriminate and try to only let Google's scrapers access the data - what's wrong with that? Scraper brand is not a protected class. Tough taters if your business model "depends" on your ability to successfully make requests to another uninvolved company's webservers.
If I were the judge, I'd throw this out and let LinkedIn/HiQ duke it out themselves - they deserve each other.
Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]
#185Earlier quoted context omitted.
Lol really? I'm not "on" your site when I browse there. I asked your server to send me some data and it did so. Its real life equivalent to social engineering. Its so far not illegal for me to ask you things and for you to disclose them to me even if you weren't supposed to. I'm allowed to lie to you even to persuade you to tell me things.
You didn't "ask my server". You used a tool to extract data from my server. It's more akin to you standing just outside my property border and using a fishing pole to pull fish from a pond that is inside my property border. You're still trespassing even if your two feet aren't physically on my land. The common legal argument (see the second link in my above comment) is that accessing a web server actually does consti…
Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]
#186Earlier quoted context omitted.
I feel like this is a really common theme I've seen several times. Something like "Music Lyric site X sues Google for embedding their lyrics in the results directly" which is funny because site X got the lyrics by scraping them from other sites. Plus Google only exists from scraping content, but I believe their TOS includes "don't scrape our content". I find it really funny that the scrapers are battling scrapers - l…
> Plus Google only exists from scraping content, but I believe their TOS includes "don't scrape our content". Yes. This is EXTREMELY frustrating. Of all companies to prevent scraping, Google is the most ironic. Especially since their goal is to organize the world's information, it shocks me that there's no way to get access to this organized information from machine to machine.
Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]
#187Earlier quoted context omitted.
I feel like this is a really common theme I've seen several times. Something like "Music Lyric site X sues Google for embedding their lyrics in the results directly" which is funny because site X got the lyrics by scraping them from other sites. Plus Google only exists from scraping content, but I believe their TOS includes "don't scrape our content". I find it really funny that the scrapers are battling scrapers - l…
Google respects robots.txt, so it's arguably not the same as scraping a website without their consent.
Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]
#188I'm reminded of Dave Chapelle's Halle Berry routine...
Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]
#189Earlier quoted context omitted.
Not allowing the CFAA to be (ab)used to attempt to make scraping illegal makes sense. However, how is it reasonable to force a web site to serve its contents to a third-party company, without being allowed to make a decision whether to serve it or not? Serving the web site costs money, and the scraper surely isn't going to generate ad income...
Isn't the issue of being selective on who can view the content? If I, random Joe User views the publicly available content you have no issue. But if someone scrapes that data them you'd want to charge them. Unless I click on the ad, the act of using your bandwidth doesn't change based on who the viewer is. You'd want to apply fees based on the future use of the data rather than on your actual costs.
I could see the hit from a scraper being heavier than that of a typical user. There's also the potential that a user is going to click an ad for any number of reasons, there isn't that likelihood the scraper will.
I'm not anti-scraping by any means, but I get the concerns.
Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]
#190This action does more than that. The court left the preliminary injunction against LinkedIn in place: "The district court granted hiQ’s motion. It ordered LinkedIn to withdraw its cease-and-desist letter, to remove any existing technical barriers to hiQ’s access to public profiles, and to refrain from putting in place any legal or technical measures with the effect of blocking hiQ’s access to public profiles." So Lin…
Not allowing the CFAA to be (ab)used to attempt to make scraping illegal makes sense. However, how is it reasonable to force a web site to serve its contents to a third-party company, without being allowed to make a decision whether to serve it or not? Serving the web site costs money, and the scraper surely isn't going to generate ad income...