Keep in mind that companies have sued for scraping not through the API, for example LinkedIn, which explicitly prevents scraping via the ToS: http://www.informationweek.com/software/social/linkedin-sues... OKCupid did a DMCA takedown for researchers releasing scraped data: https://www.engadget.com/2016/05/17/publicly-released-okcupi... Since both of these incidents, I now only scrape if it's a) through the API follow…
This post is kind of crazy, aggrandizing bad behavior and misuse of other's resources against their will. Scraping against the TOS is super bad netizen stuff, and I dont think people should be posting positive reviews of people doing this. Breaking captchas and the like is basically blackhat work and should be looked down upon, not congratulated as I see in this thread.
Web Scraping in 2016
361–370 of 402 posts
Re: Web Scraping in 2016
#362Earlier quoted context omitted.
So if instead I scrape their site (like they are scraping others) I don't have any opportunity to agree to their terms? Much like their scrapers on other sites? I'm honestly wondering about the double standard. There is a rational way to discuss morality/ethics and subsequent laws regarding most technical aspects, that often mirrors real world (read: offline/analog) scenarios. It's unfortunate that the legal system h…
>> It's unfortunate that the legal system has instead been appropriated by lawyers. omg, really? It's unfortunate that the internet has instead been appropriated by hackers. It's unfortunate that the stock market has instead been appropriated by traders. It's unfortunate that the asylum has instead been appropriated by inmates.
Very few traders went to jail after 2008. Seemingly legal (or at least not illegal). Should they have? Most bright/talented lawyers are likely working (again within the law) to get megacorps or rich people off for something poorer people would not. In our field this OP is one of the issues. What information is free and what information is not? What things I'm allowed to do offline am I allowed to do online?
I'm not proposing a solution, but any system populated by humans will be abused by some, and fought for by some idealists, all within that systems rules.
Let's take murder: I stab someone: murder. I use a broom to push a flower pot off a balcony hitting someone in the head, killing them: murder. I swat a butterfly in Beijing, causing a chain of events to a container crushing a dock worker in Rotterdam. Murder? If this extreme example comes down to intent it's thought crime, otherwise I'm playing within the rules of the system, and I just killed someone, scot-free.
While there apparently were no laws prohibiting the upsale of bad mortgages, and banks having the resources to move the market towards more and worse mortgages, that also was within the systems rules, but I personally think it's far beyond the intended use of that market, and well outside the spirit of the laws.
There's a huge difference between judicial justice and what most would agree was "justice". That's where my first comment came in. True about most systems.
Re: Web Scraping in 2016
#363Keep in mind that companies have sued for scraping not through the API, for example LinkedIn, which explicitly prevents scraping via the ToS: http://www.informationweek.com/software/social/linkedin-sues... OKCupid did a DMCA takedown for researchers releasing scraped data: https://www.engadget.com/2016/05/17/publicly-released-okcupi... Since both of these incidents, I now only scrape if it's a) through the API follow…
It would be ok if it wasn't "You can't scrap my site. Unless of course you're Google" this double standard drives me mad.
Re: Web Scraping in 2016
#364Earlier quoted context omitted.
Incidentally, this only further proves my point. If you're a big company that's retained massive law firms, you can successfully raise a fair use and implied license defense. If you're not, you can neither mount a strong offense against that defense nor mount a strong defense against Google's hypocritical offense if you find yourself on the other side. Google's primary out here is its reputation (not guarantee) for o…
You claimed "[Google] infringes on copyright as a matter of course" despite the many real civil cases (previously cited) which have found these very activities to be non-infringing. And then, strangely you claimed: >Fair use is a case-by-case basis, so you can't say that Google's infringing conduct is generally accepted to be fair use. There is so much wrong with this statement. For one, how can you call something in…
Fair use is an affirmative defense. Google admits that it copies content without legal license to do so, but claims that said copies are non-infringing under fair use exemptions. I guess you're probably correct that it's no longer appropriate to refer to Google's behavior specifically as "infringing", just "copying without authorization", which, for those of us without $5 million to commit to a legal team, means "infringing". I will try to remember the special standard of law which has been allowed to Google and refer to their copying only as "unauthorized" and not "infringing" in the future.
If you review the points summarized in the Wikipedia articles you helpfully linked, you'll see that Google's defense is mostly "Yeah, but we're Google".
In Field, "the court found that the plaintiff had granted Google an implied, nonexclusive license to display the work because of Field’s failure in using meta tags to prevent his site from being cached by Google.", i.e., because Field already knew Google existed and knew there was a standard way to prevent its access but chose not to employ it, he gave Google an implied license.
Who else does that work for? Can I send an email to Netflix and tell them "Hey, if you don't want me to copy your shows, please add this in your page's HEAD element: "? No?
I understand there are other criteria which were used to decide if Google's use was specifically infringing in addition to the implied license. Just demonstrating that Google is getting favored treatment from the judiciary that would not be available to a normal entity.
In Perfect 10 [0], the judge even explicitly indicated that he was loathe to find Google's use of thumbnails infringing because he didn't want to "impede the advance of internet technology", but that he felt the law obligated him to do so (his ruling in that matter was overturned on appeal, when the Ninth Circuit found Google's usage non-infringing). What if the defendant had been some company perceived as less technically advanced than Google? This is probably as close as you can get to an explicit statement of favoritism. The Ninth Circuit also rejected Perfect 10's claim that RAM copies were infringing (which was not the case with an unlucky non-Google company discussed further down).
What if I started indexing and rehosting thumbnails? I can assure you that I would get C&D'd almost immediately and I would be forced to shut down because I can't afford to pay lawyers for 3 years while the case works through the system (and to be honest, I'm surprised it only took 3 years). And even if I could, with a reputation less sterling than Google's, there's no reason to believe that a judge would rule in the favor of one useless guy instead of a big company. A judge would look at the case and say "Google's use was fair because it provided a public service [actually cited as part of the justification in most of your linked cases], but this guy is just using it for a few hundred people, it's definitely unfair, he owes that company more money than he'll make in his life, case dismissed".
There are many such cases on the books. I don't know if Google has a direct connection to the reptilian overlords or what, but it seems in most cases where they're not involved, the good side loses.
In Craigslist v. 3Taps, while primarily a CFAA case, 3Taps was found to be infringing copyrights by sampling Craigslist postings in order to allow its clients to plot them on a map. Being a "public service" or a "referential use" didn't matter for them. They were raked over the coals, and it's been that way with most cases.
In Ticketmaster v. RMG Technologies [1], RMG was found to infringe just by parsing a page. "Defendant's direct liability for copyright infringement is based on the automatically-created copies of ticketmaster.com webpages that are stored on Defendant's computer each time Defendant accesses ticketmaster.com. [...] Defendant contends [...] that such copies could not give rise to copyright liability because their creation constitutes fair use[.] [...] Defendant's fair use defense fails."
The case specifically discusses how, despite the precedent in Perfect 10, since the Defendant is not Google, it is bound by a site's Terms of Use and copyright law, and RAM copies, which are specifically non-infringing for Google, were infringing for RMG.
Very similar findings were made in Facebook v. Power Ventures, and the founder was left holding a bag of $3 million in personal liability.
This is a thread about the legality of HN users scraping. It seems Google is the only entity capable of making unauthorized copies and then getting courts to agree that it's fair use. For the rest of us, it's infringement, which carries stiff penalties (and this doesn't even broach the CFAA portion of the issue).
So when I say "infringing", I mean something that would be considered infringing if you aren't Google. It's apparently only infringement if the judges involved don't personally use your site and don't have to worry about personally suffering the consequences of not having access to it. :)
[0] https://www.eff.org/document/perfect-10-v-google-ninth-circu...
[1] https://scholar.google.com/scholar_case?case=147697505884223...
Re: Web Scraping in 2016
#365Earlier quoted context omitted.
To see the terms that Google thinks you have agreed to, click 'Terms' at the bottom of www.google.com If that doesn't hold up in court, in future on your first visit to Google it will simply display some text and require that you click 'I agree' to continue. Either way, it seems reasonable to me that you should agree to their terms in order to use their service.
So if instead I scrape their site (like they are scraping others) I don't have any opportunity to agree to their terms? Much like their scrapers on other sites? I'm honestly wondering about the double standard. There is a rational way to discuss morality/ethics and subsequent laws regarding most technical aspects, that often mirrors real world (read: offline/analog) scenarios. It's unfortunate that the legal system h…
Re: Web Scraping in 2016
#366Earlier quoted context omitted.
There are hundreds of paid services that scrape Google heavily (search engine ranking trackers). How are they legal?
They probably doing it from country where it's legal. In most countries there is no law that would be applicable in this case.
Re: Web Scraping in 2016
#367Earlier quoted context omitted.
>How can TOS have legal power for the case scraping? A website is a public property. If I'm visiting it without logging in, I don't have a chance to accept TOS. This is called "clickwrap". There is usually a notice in the footer of each page that says something like "By using this site, you agree to our Terms of Service." Typically, this kind of notice has been held enforceable. More recently, judges have been demand…
> Thus, if you take a photograph of a building built in 1991 and the year is not yet 2111, there is a chance that the architect can claim infringement. The architect can claim infringement all they want, they don't have a case. From https://www.law.cornell.edu/uscode/text/17/120 : The copyright in an architectural work that has been constructed does not include the right to prevent the making, distributing, or public…
This is an important caveat to architectural copyright, however, so thanks for clarifying.
Re: Web Scraping in 2016
#368Earlier quoted context omitted.
Which TOS? I might have accepted terms when I created a Google Account but in no way do I agree to a TOS by visiting a URL.
To see the terms that Google thinks you have agreed to, click 'Terms' at the bottom of www.google.com If that doesn't hold up in court, in future on your first visit to Google it will simply display some text and require that you click 'I agree' to continue. Either way, it seems reasonable to me that you should agree to their terms in order to use their service.
This statement is demonstrably false, as shown by all the places in the world where this type of TOS-nonsense actually does not hold up in court.
And in the USA, it's (as usual) even slightly more absurd: The only reason it does hold up in court is because Google can afford justice.
Re: Web Scraping in 2016
#369Earlier quoted context omitted.
>> It's unfortunate that the legal system has instead been appropriated by lawyers. omg, really? It's unfortunate that the internet has instead been appropriated by hackers. It's unfortunate that the stock market has instead been appropriated by traders. It's unfortunate that the asylum has instead been appropriated by inmates.
To some extent, yes. When people spend enough time in their given field to know the ins and outs, those less scrupulous tend to bend the rules more and more. While not _strictly_ against the rules it often ends up going against the spirit at the base of the industry. Very few traders went to jail after 2008. Seemingly legal (or at least not illegal). Should they have? Most bright/talented lawyers are likely working (…
Re: Web Scraping in 2016
#370Earlier quoted context omitted.
Given that FB's bot identifies itself, no, eventually some websites will present og: markup only to FB's bot.
Couldn't you just spoof the user agent though? Or is there some other mechanism that you can use to verify the bot is Facebook's?