Live data from Hacker News

Congrats! Web scraping is legal! (US precedent)

parsers.me

271–280 of 409 posts

Re: Congrats! Web scraping is legal! (US precedent)

#272
post #75

Anyone ever read any stories about the concept of "agents" where everyone had their own computer agents that did all the work and talked to other agents do all the steps like book your flights, tickets, order your food, collect/collate data for your purposes, etc? We need to make the internet "Agent" friendly.. we should stop assuming the end user (end human?) will ever see any webpage on the internet.

Having an agent proxy for you won't last long. If that agent has connections to your friends' agents, your place of work, etc., then the protection provided by the proxy whittles away quickly.

To find you, all anyone would need is a photo of you, match it to an existing photo tagged as your agent, and your agent's online presence is clearly identified.

You could rotate to a new agent every time you go online, or randomly while online, but forget about making purchases, maintaining friendships or building "cred".

The only sure way to separate you, the person, from your online presence is to not have an online presence. For the company in question above, the only way to isolate yourself from their "running to tell Mom" business model is to simply not put your resume up on LinkedIn.

This ruling is troubling in many ways. HiQ can scrape the web, and now, apparently, so can every other personal information brokers out there. (Ref: https://en.wikipedia.org/wiki/Information_broker, which naively suggests this information is only important to advertisers; it is also of interest to governments, police, employers, even the parents of the S.O. you intend on marrying, or the S.O. themselves.)

Is there a company that will track your online posts while you are supposed to be working, and "run and tell Mom" that you were doing something other than work? As an employer, would you pay for such a service?

I'm not even convinced that air-gapping your person from the internet would work in this new world. A lack of online presence, while not damming in itself (yet), could indicate "something to hide."

Re: Congrats! Web scraping is legal! (US precedent)

#273
post #7

> Now many site owners are trying to put technical obstacles to competitors who completely copy their information that is not protected by copyright. For example, ticket prices, product lots, open user profiles, and so on. Some sites consider this information “their own”, and consider web scraping as “theft”. Legally, this is not the case, which is now officially enshrined in the US. Does this mean we can now scrape…

Yes you can scrape them, no you cannot repubilsh them. Everything you listed is protected by copyright. You cannot infringe on copyrights because of this ruling. >hiQ argued that LinkedIn’s technical measures to block web scraping interfere with hiQ’s contracts with its own customers who rely on this data. In legal jargon, this is called” malicious interference with a contract”, which is prohibited by American law Do…

No, because what one side of a case argues is not the law. What judges decide is the law.

Re: Congrats! Web scraping is legal! (US precedent)

#274

Earlier quoted context omitted.

> Lots of sites have ToS preventing such things, are those legally void now? Are captchas on public pages illegal, even if you request the page 8000 times in a second? ToS are subservient to the law; you can (probably) terminate a service account from a user that breaks your ToS, but if the user does not have a service account (as is the case for HiQ, it doesn't seem they were using accounts for it), then your ToS do…

> but if the user does not have a service account (as is the case for HiQ, it doesn't seem they were using accounts for it), then your ToS does not apply, since you've technically not entered a binding legal contract with them. Are you sure about this? I am not a lawyer, but I believe that the Terms of Service applies to all users, not just those that explicitly set up a user account. I have interpreted the LinkedIn…

> Are you sure about this? I am not a lawyer, but I believe that the Terms of Service applies to all users, not just those that explicitly set up a user account.

Typically you'll see TOS say something along the lines of "by continuing to access this site you agree..." or "if you do not agree with these terms you may not access this site..."

Whether that's enough to create a binding contract depends on the jurisdiction and who you ask.

Re: Congrats! Web scraping is legal! (US precedent)

#275
I'd been following this for sometime, happy that the appeals court upheld common sense reasoning. I always found it some what ridiculous that being able to crawl publicly available data was such a grey area. As long as you're being a responsible crawler there is no downside to the company whose data is being scraped other than in the anti competitive sense.

Re: Congrats! Web scraping is legal! (US precedent)

#276

So many ideas start to come to mind if scraping is legal. Can we start to scrape Google Search in order to bootstrap building an alternative to Google Search? Search is a really hard problem (that somebody should tackle), but if we can leverage what Google has already scraped from the web and associated with popular search terms, we can use that to help train and validate our search model. Can we scrape Reddit, Twitt…

> Can we scrape Reddit, Twitter, or Facebook in order to stand up a competing service that strips out all the ads?

How are you going to pay for it? Subscription model doesn't work for search/social networks.

Re: Congrats! Web scraping is legal! (US precedent)

#277

Earlier quoted context omitted.

It seems absurd if the 'interference' only directly affects their own property. Like, if my neighbors start monetizing livestreaming my backyard, suddenly I can't put up a fence? Except worse because in actuality, this third-party contract is costing them money through server load and bandwidth.

Your analogy doesn't hold. Your backyard is private property. The data that LinkedIn publishes is intended for the public. That's why Google can index the pages and give you results from LinkedIn.

Can someone taking down a open source project, like the leftpad debacle, be sued for tort?

Re: Congrats! Web scraping is legal! (US precedent)

#278
post #243

Earlier quoted context omitted.

Why can't I have terms on my website that say how you can use my information? Examples where this is allowed: - Images/media (Creative commons) - Code (Open source licenses) You say it isn't allowed for: - Personal data Unless I'm misunderstanding your philosophy (which seems to say copyright is OK, but public information must be public to all): You believe that it's morally OK for me to prevent a company selling my…

> Why can't I have terms on my website that say how you can use my information? Neither Creative Commons nor copyleft (nor copyright in general!) can assert anything about private use . IP rights are commercial rights; they affect sellers of your IP. They don’t affect end-consumers of your IP. Note that even the GPL can’t force someone to publish the source of their GPLed-library-containing program, if they never pub…

Right, sure, I hadn't really thought about how copyright isn't really enforceable against individuals. That's very interesting.

However, I'm really not sure how this is relevant to your moral stance on commercial use of "public" personal information.

Why do you believe it's reasonable to prevent unauthorized commercial exploitation of creative works, but not to prevent unauthorized commercial exploitation of personal information.

The former simply affects the small percentage of people who sell their works.

The latter affects the vast majority of the population who receive targeted spam, have their information collated and sold for profiling, are victims of identity fraud when those databases are inevitably leaked, etc.

For what it's worth, as I mentioned in my first comment, the GDPR absolutely gives me rights to control how my personal information is used. And the GDPR has a near total exemption for individual use.

What benefits do you see of commercial use against the wishes of the person that published it that outweigh the risks? (making money isn't a benefit)

Re: Congrats! Web scraping is legal! (US precedent)

#279

This only affects the ninth circuit—which includes the tech hubs San Francisco, Seattle, LA, and Portland. It would only apply to the rest of the country if the Supreme Court affirmed it. Even then, a well-funded company or zealous prosecutor could say that it doesn’t apply in your case because of some technicality. In that case you would need hundreds of thousands or millions of dollars and a few years to litigate t…

Unfortunately, the battle isn’t over yet.

LinkedIn does plan on bringing this to the Supreme Court: https://www.law360.com/articles/1237505/linkedin-will-go-to-...

Re: Congrats! Web scraping is legal! (US precedent)

#280

Earlier quoted context omitted.

> Lots of sites have ToS preventing such things, are those legally void now? Are captchas on public pages illegal, even if you request the page 8000 times in a second? ToS are subservient to the law; you can (probably) terminate a service account from a user that breaks your ToS, but if the user does not have a service account (as is the case for HiQ, it doesn't seem they were using accounts for it), then your ToS do…

> but if the user does not have a service account (as is the case for HiQ, it doesn't seem they were using accounts for it), then your ToS does not apply, since you've technically not entered a binding legal contract with them. Are you sure about this? I am not a lawyer, but I believe that the Terms of Service applies to all users, not just those that explicitly set up a user account. I have interpreted the LinkedIn…

Also not a lawyer, but you cannot force me to accept your terms of service. Contract law requires both parties agree to enter it.

When you create an account, etc., you are agreeing to those terms. If I browse a public webpage that just has a terms of service link on the bottom of it, I've not agreed to anything.

Post reply on HN