Does this mean Google can't stop me from scraping their results anymore? If I do it from an US ip address of course.
Congrats! Web scraping is legal! (US precedent)
351–360 of 409 posts
Re: Congrats! Web scraping is legal! (US precedent)
#352Re: Congrats! Web scraping is legal! (US precedent)
#353The toxicity towards web-scraping is really what makes me lose hope in the current web. People want their data to be public and all of the benefits that comes with public data but then they want to chose who gets to see it - it's a complete and utter paradox. This precedent doesn't really mean much but is definitely step in the right direction.
> People want their data to be public and all of the benefits that comes with public data but then they want to chose who gets to see it - it's a complete and utter paradox. That's a complete misconception. Of course you can find manufacture inconsistent ideologies if you combine ideas from different people, but I think you'd have a difficult time finding one person who believes what you just described. What I want i…
Re: Congrats! Web scraping is legal! (US precedent)
#354Linkedin is taking this to the Supreme Court: https://www.law360.com/articles/1237505/linkedin-will-go-to-... No ultimate decision was ever made, and no, this doesn't make web scraping 100% legal. Wake me up when there's a new announcement because anyone interested in this already know this old news.
This is a really big deal. Currently (IMHO) the US Supreme Court is a wholly-owned subsidiary of multinational corporations due to the shenanigans that happened with Obama, McConnell and Garland, so will likely side with LinkedIn since it's the larger corporation: https://www.npr.org/2018/06/29/624467256/what-happened-with-... I feel like siding with LinkedIn here would open up the web to extortion though, like troll…
Re: Congrats! Web scraping is legal! (US precedent)
#355Earlier quoted context omitted.
> as an owner you are legally allowed to disallow service to anyone without having to provide a reason as a website owner you absolutely have the right to limit who does or doesn't visit your site, but that's up to you to enforce. Scraping is literally equivalent to reading a giant public billboard and writing what it says somewhere else. How could that be illegal?
> Scraping is literally equivalent to reading a giant public billboard and writing what it says somewhere else. How could that be illegal? If you treat the website you're visiting as a privately owned business then you're trespassing on private property by scraping their site, assuming that site is trying to prevent that behavior. If you don't treat the website as a privately owned business, then what do you classify…
As long as there are big glass windows, the patrons and the information gained from looking inside are not subject to privacy, and are thus freely accessible to anyone walking by.
What you're not entitled to is the backend workings of the website - the application/code/databases, and credentials. It would be akin to going into a restaurant and forcing yourself into the food prep area and stealing the owner's keys, and contaminating the food that patrons eat.
Scraping too often would be something akin to setting up a camera tripod right in the building entrance, hindering the influx of new customers. While not explicitly illegal, it does hinder the operation of the business, and you will probably be physically removed at some point (this is banning abusive bots, etc.)
Re: Congrats! Web scraping is legal! (US precedent)
#356Does this mean someone can scrape a restaurant directory website and create their own website with this same data?
Re: Congrats! Web scraping is legal! (US precedent)
#357I guess this means the web is lawfully seen as belonging to the public which seems both good and bad, because while it might be publicly accessible as a whole, each individual site is privately owned. I'm pretty sure with any brick and mortar business, as a business owner you are legally allowed to disallow service to anyone without having to provide a reason. That would be the equivalent of allowing to block bots /…
It depends, you aren't allowed to restrict access for privileged reasons (i.e. banning all black people from your business is a clear case of discrimination and is specifically illegal), but otherwise you have a lot of levity to refuse persons entry - but that "without having to provide a reason" is just apparent because it is very infrequently challenged - if someone feels they were wrongly refused service they absolutely have the right to sue and require you to state why that ban took place, but, as your business is owned by you a ban made by you is legal until overturned, either early (as in a preliminary injection) or as a result (after the ruling on a case).
It's hard to pull real world analogies for this case, as well, since Amazon is never worried about other customers being dissuaded from a purchase because I'm shopping in my boxers, or am shouting and creating an unsafe environment - so there's a limit to how well you could tune a real world example to this scenario... but I might give it a try with this:
I run an absolutely 'grammable business, there's bubbly pink champagne and gold everywhere - including the stoop. I like to use that fancy stoop to attract new customers but I've been a bit concerned lately that a city tour guide has regularly been bringing around their tour and they're taking selfies of my business from across the street - that tour guide is profiting from the decoration I've put into my stoop and some of those tourists might've been willing to pay the modest 300$ fee to walk in and take selfies inside! So I try to get the city to restrict the tour guide from being able to lead his tourists to take selfies from across the street. It's very possible I could remove the glitzy decorations from my stoop but that would impair my advertising, which I don't want to do - instead I want to legally restrict people from using my stoop in a manner I did not intend.
I think that's pretty apt in a few ways but it also misses the fact that LinkedIn is technically paying a bit of money every time someone loads a page - and if the tour guide was noticeably damaging my stoop I would be able to sue him for damages and probably win, but the cost of that page load is so utterly negligible that maybe we can ignore it.
Re: Congrats! Web scraping is legal! (US precedent)
#358Earlier quoted context omitted.
> People want their data to be public and all of the benefits that comes with public data but then they want to chose who gets to see it - it's a complete and utter paradox. That's a complete misconception. Of course you can find manufacture inconsistent ideologies if you combine ideas from different people, but I think you'd have a difficult time finding one person who believes what you just described. What I want i…
Making organization membership public would trample on personal privacy quite effectively in some respects, such as with disease support groups or PACs; medical privacy is taken seriously, but is there such a thing as political affiliation being private? Is it a violation of someone's privacy to reveal they give to the ACLU?
In a world where organizations are radically transparent and individuals have radical privacy, before you join the disease support group or donate to the ACLU, you already know the organization's records about that transaction will be public information.
You can restrict your activities to organizations that do not keep personally identifiable information. Or you could join the support group with a pseudonymous identity or donate cryptocurrency to the ACLU.
Re: Congrats! Web scraping is legal! (US precedent)
#359Earlier quoted context omitted.
Wouldn't the solution be to offer a streamlined download (maybe even as a torrent if you're worried about bandwidth) of all the data then?
For what purpose? That’s like suggesting that if people keep jumping your fence and trampling your roses because it’s a shortcut to a public park (in this case, the county records office) that already has public access roads, that you should be obliged to build a sidewalk through your garden, at your own expense, when the real answer should be that the public road should be improved.
The ideal answer and the efficient answer are not usually the same.
Re: Congrats! Web scraping is legal! (US precedent)
#360So, if I'm reading this right:
If it's not behind a login, it's free game for scraping. CFAA only explicitly applies once you (the service provider) enters into an explicit relationship of use with someone else.
I.e. If LinkedIn made an account necessary to even browse the currently "public" data, and set limits as a condition of being a user that would be fine, and may give them a leg to stand on for future claims of misuse of their systems since the behavior is already covered by a default terms of use in creating an account.
In terms of making currently public data more difficult to access in response to scraping, I foresee business evolving even more in the direction through which everything gets locked behind an initial contract step to ensure that people can reserve the right to refuse access.
I'm very wary of the whole malicious interference with a contract reasoning that hiQ employed. Just because you're relying on a public data source to be accessible in a convenient way to you to deliver a service to an unrelated third-party should not obligate the party of the public data source to deliver data in a convenient way to you. That type of thing would lead to ridiculousness like being able to sue the makers of the Yellow/White Pages for instance for changing things up.
Further, I still see grounds for remedy to LinkedIn in that they do still have a right to refuse service to anyone, regardless of outstanding contacts entered into by that person. The only leg to stand on that I can see for hiQ is if the court limited the precedent being set in the event that LinkedIn identified them, notified them and requested them to cease and desist, hiQ refused (or hiQ offered to enter into an alternate business arrangement and LinkedIn refused to accomodate), and then LinkedIn implemented the interfering measure. In that circumstance, I can see a malicious interference maybe being upheld, but I still have very little sympathy for hiQ because whether the pages are public or not, they are imposing a significant cost on LinkedIn to serve that data if his is systematically enumerating the entire public LinkedIn dataset.
To me, hiQ is a guy OCR'ing the White Pages/recording a storefront 24/7. The details of how the Internet operates changes that a bit, but nevertheless, it is there.
It is a fundamental continuation of a trend I've noticed in Tech in that it is a field where practitioners/the businesses practitioners write code for for whatever reason assume that the platform running/servicing code they write fundamentally "belongs" to them. That any data their program can get access to is free-game.
I don't know where the disconnect is, or if maybe I'm just a freak in that I treat someone else's hardware as if I were entering their home. Even with the code I write.
You don't just waltz into someone else's space and start using all the amenities without asking. It's rude, and grounds to have admittance refused in the future. You don't do things to the system without asking.
It's like... Imagine most people are blind. You're a salesman walking into their space talking a good story or providing some service while completely ransacking their domicile for every shred of info you can possibly take with you. Filming the inside, the layout, rifling through mail and rolodexes, so on and so forth, while all they are aware of is explicitly what you tell them is going on.
As one who writes software, and automates tasks with computers, I feel it to be our duty to not facilitate that sort of behavior. I just wish I could figure out a way to spread that ethos and keep the bills paid.
Otherwise I fear we'll lose any benefit of an assumption of integrity we've managed to build up. Hell, maybe it's too late for that given the way things are going.