Live data from Hacker News

Congrats! Web scraping is legal! (US precedent)

parsers.me

371–380 of 409 posts

Re: Congrats! Web scraping is legal! (US precedent)

#371

Earlier quoted context omitted.

I'm going to assume you're asking in good faith and try to address the confusion here. The human does get rights, the organization doesn't. In some cases, believing that humans have rights and believing that organizations have rights might lead one to the same action. In those cases, I'd take the action. I wouldn't want to violate a human's rights out of some vindictive dislike of organizations: that's not the point.…

> The human does get rights, the organization doesn't. Organisations are (usually) legal persons, too; they just have fewer responsibilities, get fewer rights in exchange.

That's how things are, not how they should be.

Re: Congrats! Web scraping is legal! (US precedent)

#372

Earlier quoted context omitted.

Let's say that individual humans have the right to keep secrets. Let's also say that they have the right to keep secrets with their associates, and to tell them to who they please. Now, doesn't that make it legal for a group of people to keep secrets about you ? What about selling them? I just don't see what doing away with the legal fiction of corporate personage would do about Facebook.

"Now, doesn't that make it legal for a group of people to keep secrets about you?" A group of people sure, but corporations are not people.

Corporations are groups of people.

Re: Congrats! Web scraping is legal! (US precedent)

#373

Earlier quoted context omitted.

"Now, doesn't that make it legal for a group of people to keep secrets about you?" A group of people sure, but corporations are not people.

Corporations are groups of people.

Corporations are a power of attorney document. They are golems that sometimes act on behalf of people. They are not people.

Re: Congrats! Web scraping is legal! (US precedent)

#374
post #343

I strongly disagree that not allowing scraping protection on social network is a good thing (reasons mostly from this thread https://news.ycombinator.com/item?id=22182144 ) What I would do on Linkedin side is: 1. Split public setting into 2 settings: public for everyone including scraping, and second choice that make it public for Linkedin users (aka banning scraping [1]) 2. Everyone on Linkedin who have current sett…

Scrapers can just create a user account then. They still seem to be covered by this injunction

Re: Congrats! Web scraping is legal! (US precedent)

#375

Earlier quoted context omitted.

> but if the user does not have a service account (as is the case for HiQ, it doesn't seem they were using accounts for it), then your ToS does not apply, since you've technically not entered a binding legal contract with them. Are you sure about this? I am not a lawyer, but I believe that the Terms of Service applies to all users, not just those that explicitly set up a user account. I have interpreted the LinkedIn…

> Are you sure about this? I am not a lawyer, but I believe that the Terms of Service applies to all users, not just those that explicitly set up a user account. Typically you'll see TOS say something along the lines of "by continuing to access this site you agree..." or "if you do not agree with these terms you may not access this site..." Whether that's enough to create a binding contract depends on the jurisdictio…

It can also depend on the terms themselves. I can put "by using this site you agree to bake me a chocolate cake" on my website all day, but that doesn't mean I will be able to force you to bake me a chocolate cake.

Re: Congrats! Web scraping is legal! (US precedent)

#376
post #174

Earlier quoted context omitted.

> People want their data to be public and all of the benefits that comes with public data but then they want to chose who gets to see it - it's a complete and utter paradox. That's a complete misconception. Of course you can find manufacture inconsistent ideologies if you combine ideas from different people, but I think you'd have a difficult time finding one person who believes what you just described. What I want i…

I'm interested in hearing your take on "organizational transparency". Like please push the concept / idea to its 'full' realization and tell me that picture, even if it implies a little bit of "sci-fi"¹. Digging this because I think that domain / paradigm will see unparalleled evolution in the next few decades. [1]: I mean, don't stop at current law / values / behaviors; like people from the 1940s wouldn't have dared…

I don't think that looking too far ahead is useful: this is just a matter of pragmatics. Revolutionary change in a peaceful society happens via a long sequence of small, incremental changes, and that's a good thing, because you get to see how each of the changes plays out. I think the best sci-fi persuades you that it's looking at the distant future when in fact it's only using the future as a foil to provide deep insight into the present.

The short-term, the small, incremental changes I'd like to see are:

1. Reversal of the default privacy setting of government docs. Instead of documents being default-private and citizens having to make FOIA requests to make those documents public, documents should be default-public, and government workers should have to apply through and adversarial system (similar to courts) to classify documents, proving to a court why the document needs to be classified.

2. Classified documents should have a short (1 year max) timeframe after which they are declassified, or government workers should have to reapply to justify why the documents need to remain classified.

3. Political party documents should be public, without any provision for classifying them.

4. Tax-exempt organization documents should be public, without any provision for classifying them.

5. IPO'ed organization documents should be public, without any provision for classifying them.

6. Body cams on all police and military while on duty (when they are acting on behalf of an organization). 1 and 2 would apply to the footage from these cams as well.

7. Exceptions to 1-6 should be made for the personally-identifiable information of people who are not in the organization.

8. Organizations should be required to maintain a list of all the personally-identifiable information they have on a person (including employees), and provide that data to that person on demand by that person or their legal guardian, as well as a list of all people with whom that data has been shared, and be required to delete that information upon request by that person or their legal guardian.

9. Research which receives public funding should be forced to publish its results publicly.

10. All software which receives public funding should be forced to publish its source publicly.

11. Government documents should be published in open-source formats suitable for computer analysis (i.e. CSV, text, or some XML format--no PDFs).

Re: Congrats! Web scraping is legal! (US precedent)

#377

So many ideas start to come to mind if scraping is legal. Can we start to scrape Google Search in order to bootstrap building an alternative to Google Search? Search is a really hard problem (that somebody should tackle), but if we can leverage what Google has already scraped from the web and associated with popular search terms, we can use that to help train and validate our search model. Can we scrape Reddit, Twitt…

> Can we finally scrape and get rid of IMDB? I'd love to put all of their content on a wiki and be done with it. Just because you can scrape the content legally does not mean you can also republish it on your own website.

> Just because you can scrape the content legally does not mean you can also republish it on your own website.

Except IMDB copied all of its data by scraping publicly available data posted to Usenet back in the day. And they still rely on volunteer contributions. [1]

[1] https://en.wikipedia.org/wiki/IMDb#History

Re: Congrats! Web scraping is legal! (US precedent)

#378
post #163

Earlier quoted context omitted.

> People want their data to be public and all of the benefits that comes with public data but then they want to chose who gets to see it - it's a complete and utter paradox. That's a complete misconception. Of course you can find manufacture inconsistent ideologies if you combine ideas from different people, but I think you'd have a difficult time finding one person who believes what you just described. What I want i…

Making organization membership public would trample on personal privacy quite effectively in some respects, such as with disease support groups or PACs; medical privacy is taken seriously, but is there such a thing as political affiliation being private? Is it a violation of someone's privacy to reveal they give to the ACLU?

The entire point of organizational transparency is to prevent organizations from trampling the rights of individuals, so in cases where organizational transparency would trample the rights of individuals, the rights of individuals supersedes the need for organizational transparency.

Re: Congrats! Web scraping is legal! (US precedent)

#379

Earlier quoted context omitted.

> Lots of sites have ToS preventing such things, are those legally void now? Are captchas on public pages illegal, even if you request the page 8000 times in a second? ToS are subservient to the law; you can (probably) terminate a service account from a user that breaks your ToS, but if the user does not have a service account (as is the case for HiQ, it doesn't seem they were using accounts for it), then your ToS do…

> but if the user does not have a service account (as is the case for HiQ, it doesn't seem they were using accounts for it), then your ToS does not apply, since you've technically not entered a binding legal contract with them. Are you sure about this? I am not a lawyer, but I believe that the Terms of Service applies to all users, not just those that explicitly set up a user account. I have interpreted the LinkedIn…

I am a lawyer, and there isn't really an easy answer to these questions.

TOS are a lot like EULAs. If they look like contracts of adhesion, then they're going to get more scrutiny and skepticism. The TOS that you claim applies even to every single random visitor to your site where they do not in fact affirmatively agree to the terms is potentially going to look more like a contract of adhesion. That's a lot harder to enforce.

If they are used more for CYA so that you can ban undesirable accounts from your website which people explicitly agreed to when they signed up for it, or so that you can just up and alter your entire business model without having to give all of your customers refunds, then they're easier to defend.

Just my general opinion, of course. Every jurisdiction is different.

Re: Congrats! Web scraping is legal! (US precedent)

#380

Earlier quoted context omitted.

> The human does get rights, the organization doesn't. Organisations are (usually) legal persons, too; they just have fewer responsibilities, get fewer rights in exchange.

That's how things are , not how they should be .

No rights, full liability would be a bad deal too.
Post reply on HN