Live data from Hacker News

Congrats! Web scraping is legal! (US precedent)

parsers.me

91–100 of 409 posts

Re: Congrats! Web scraping is legal! (US precedent)

#91
post #11

The toxicity towards web-scraping is really what makes me lose hope in the current web. People want their data to be public and all of the benefits that comes with public data but then they want to chose who gets to see it - it's a complete and utter paradox. This precedent doesn't really mean much but is definitely step in the right direction.

One of my clients is involved in property tax collection and reporting. Property Tax records are public info, and their website allows looking up the records for any property without a login. However, the data behind this website it the _source_ of the public records, and not the public records themselves (which would be local government databases). For years now we've been in an arms race with someone using a botnet…

Maybe you could contact the scrapper? Just post magnet links on the site that allows them to get nicely formatted dump of what they want.

Re: Congrats! Web scraping is legal! (US precedent)

#93
post #11

The toxicity towards web-scraping is really what makes me lose hope in the current web. People want their data to be public and all of the benefits that comes with public data but then they want to chose who gets to see it - it's a complete and utter paradox. This precedent doesn't really mean much but is definitely step in the right direction.

I can understand why some do not want scrapers - increased traffic (with practically zero benefits to the owners) is one obvious reason. (Some people will then say "But why not just offer APIs", but that's a lot of extra work and maintenance). It's like with instagram and other social media platforms. The content creators put in the hard work, while the leeches are stealing content for their own benefit, giving zero…

Instagram is not the content creator tho

Re: Congrats! Web scraping is legal! (US precedent)

#94
One question to those who dislike web scraping as they deem it infringes copyright laws.

Given that Google scrapes LinkedIn public profiles and add data from it to it’d index and when presenting search results, is it then not discrimination that Microsoft tries to block HiQ and not google?

Re: Congrats! Web scraping is legal! (US precedent)

#95

This only affects the ninth circuit—which includes the tech hubs San Francisco, Seattle, LA, and Portland. It would only apply to the rest of the country if the Supreme Court affirmed it. Even then, a well-funded company or zealous prosecutor could say that it doesn’t apply in your case because of some technicality. In that case you would need hundreds of thousands or millions of dollars and a few years to litigate t…

> This only affects the ninth circuit—which includes the tech hubs San Francisco, Seattle, LA, and Portland. It is only binding precedent in the Ninth Circuit, it is less accurate to say it only effects the Ninth Circuit, since decisions have effect other than as binding precedent.

Circuits can and do disagree.

Re: Congrats! Web scraping is legal! (US precedent)

#97
post #11

The toxicity towards web-scraping is really what makes me lose hope in the current web. People want their data to be public and all of the benefits that comes with public data but then they want to chose who gets to see it - it's a complete and utter paradox. This precedent doesn't really mean much but is definitely step in the right direction.

One of my clients is involved in property tax collection and reporting. Property Tax records are public info, and their website allows looking up the records for any property without a login. However, the data behind this website it the _source_ of the public records, and not the public records themselves (which would be local government databases). For years now we've been in an arms race with someone using a botnet…

At work we have all of our data available publicly as easy to parse XML files, but no matter what we do the bot owner's refuse to use it. They'd rather hammer our search engine with sequential searches instead.

Re: Congrats! Web scraping is legal! (US precedent)

#98

I ran a company based on web scraping for a couple years and I heard never ending comments about how what we were doing was illegal. Thank god that conversation is over.

You are abusing other people’s websites without intention to buy something. It’s a lot like stealing.

So... Google are the world's most successful criminal syndicate?

Re: Congrats! Web scraping is legal! (US precedent)

#99
post #35
post #22

Earlier quoted context omitted.

See that's where I have problem with this. Isn't data just _data_? Lets draw some pararells to real life. If I go to public space like town square - can't I take pictures, notes and records then go home and draw my analytics from it? What if I read something in a book I bought, can't I quote it? Same thing should be with web resources even if they are creative - as long as I don't publish them I should be able to scr…

> Isn't data just data ? No. At the risk of just repeating the comment you didn't understand, creative works are not "just data" - they are copyrightable works that the owner has control over who can use them, not just for profit, but for any reason with few exceptions. You don't just get to drop someone else's work product into your algorithm without their permission.

You can't take something copyrighted by someone else and re-distribute it without their permission. However, I suspect you can capture it freely if you don't re-distribute it.

Re: Congrats! Web scraping is legal! (US precedent)

#100
post #9

Earlier quoted context omitted.

Youtube videos are definitely protected by copyright, though.

In theory, right? See the South Park WWITB issue. I believe South Park used a videoclip from youtube, and Youtube’s ContentID system removed the video South Park had used, because Youtube considered it a violation of South Park’s copyright.

Just because YouTube gets it wrong doesn't mean it's just theory. YouTube is not the only site that has automated content scanning for copyright violations. Getty and other photo sites have gotten this wrong in the same way by sending C&D letters for violations to the actual copyright holders.
Post reply on HN