Earlier quoted context omitted.
He should open up a Patreon, tip jar, something to get that funded. Could also delay results, offer reduced temporal precision and other things to differentiate use cases.
how long are you allowed to delay results, I mean not serving results is just delaying them forever but that's out. Can I delay serving results longer than chromium's default timeout?
9th Circuit holds that scraping a public website does not violate the CFAA [pdf]
211–220 of 293 posts
Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]
#212Earlier quoted context omitted.
A website isn't a person and the law doesn't expect them to act as such. It's a tool. Making a 'request' to a web server is more like turning the knob on a door: maybe the owner installed a lock, or maybe it just opens without there being a lock. But even if there isn't a lock, the law doesn't absolve you of trespassing against the door's owner just because the door itself didn't have the sentience to refuse your req…
But it's a door handle that is MEANT to be turned by the public at large. It's like putting a big "Order Inside" sign above the door to a restaurant and being surprised when people try to gain entry. You also never entered the server. The server got your request and served something back you to. You did not go inside the house and read the contents of a book on the shelf, it was read aloud to you while you are still…
By this reasoning it's impossible ever to hack anything. Even breaking password controls or cryptography is still just sending the server a request and getting something served back.
Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]
#213Earlier quoted context omitted.
>When it comes to physical properties there's a huge difference between reading a banner posted in a street and entering the property to read some secret data: you have to be in different locations. That's why your analogy is completely faulty. At no point is accessing a web server similar in any matter to reading words off of a banner posted in a street. You cannot use a faulty analogy of your own to describe why my…
The "don't walk into someone else's house" rule applies to ALL houses everywhere. You are explicitly forbidden to enter a house unless explicitly authorized. When it comes to website, there are billions of domains in the planet, each one has multiple internal URLs, ranging from tens to several million. You can't expect everyone to have common knowledge about every domain and link. It is beyond ridiculous to compare t…
It's true that there's a presumption that sites that are accessible by the public are open for access to the public. But a lack of technical restriction is not an invitation. If a reasonable person would conclude that your access is not welcome then your access is also illegal. this is the crux of why so much of security research is on precarious legal footing. If you find an unsecured mongoDB database with a name like "customer_data" and you download the contents you are 100% breaking the law.
Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]
#214Earlier quoted context omitted.
Ugh, yeah, the more I think about this ruling, the less I like it. It's actually pretty insane to force a site to serve content. I think both parties are in the wrong here - HiQ for assuming they're entitled to receive a response from LinkedIn's webservers, and LinkedIn for abusing the CFAA to try to deny service rather than figure out a technical solution to their business problem. In my view: * The data is public,…
From reading the opinion, I think the argument goes something like this: > First, LinkedIn does not contest hiQ’s evidence that contracts exist between hiQ and some customers, including eBay, Capital One, and GoDaddy > Second, hiQ will likely be able to establish that LinkedIn knew of hiQ’s scraping activity and products for some time. LinkedIn began sending representatives to hiQ’s Elevate conferences in October 201…
Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]
#215Earlier quoted context omitted.
From reading the opinion, I think the argument goes something like this: > First, LinkedIn does not contest hiQ’s evidence that contracts exist between hiQ and some customers, including eBay, Capital One, and GoDaddy > Second, hiQ will likely be able to establish that LinkedIn knew of hiQ’s scraping activity and products for some time. LinkedIn began sending representatives to hiQ’s Elevate conferences in October 201…
Why does there need to be a legitimate business purpose? What about freedom of speech? It's my website and I'll publish what I want to.
Edit: What I mean is that freedom of speech is not the same as freedom of censoring.
Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]
#216Earlier quoted context omitted.
Why does there need to be a legitimate business purpose? What about freedom of speech? It's my website and I'll publish what I want to.
Eh, I think you got this backwards. If you really want to talk about this in terms of freedom of speech, LinkedIn is in the act of censoring? Edit: What I mean is that freedom of speech is not the same as freedom of censoring.
Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]
#217Earlier quoted context omitted.
Pictures and such are covered by copyright, but mere facts are not (at least in the US): https://en.wikipedia.org/wiki/Feist_Publications,_Inc.,_v._R... .
Collections of facts are copyrightable. Which is why you can go out and make a map of your local area but you may not copy the data from google maps. You may end up with the exact same data and that is ok because you both copied the same facts but if there is a mistake on google maps (Perhaps placed as a trap) then you can be caught if your map has the same mistake.
As for trap streets, they don't help as much as you think; as we've seen, simply showing that copying occurred is not enough. There was a theory that the trap streets - since they were invented, not facts - could themselves be copyrighted, but in Alexandria Drafting Co. v. Amsterdam, the courts said that copying a few trap streets among a bunch of facts was too minimal to be considered infringement.
Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]
#218Earlier quoted context omitted.
From reading the opinion, I think the argument goes something like this: > First, LinkedIn does not contest hiQ’s evidence that contracts exist between hiQ and some customers, including eBay, Capital One, and GoDaddy > Second, hiQ will likely be able to establish that LinkedIn knew of hiQ’s scraping activity and products for some time. LinkedIn began sending representatives to hiQ’s Elevate conferences in October 201…
That’s quite ... crazy. Be restaurant. Be on Deliveroo. Be getting low margins because of high fees. So basically you can’t decide not to use Deliveroo any more, to improve margina (“secure an exonomic advantage”). I mean, you can cancel Deliveroo, but only as long as you’re not “inducing a breach of their contract”. So only a matter of time before Deliveroo writes a contract “we’re obligated to deliver food for you…
I don't know Deliveroo, but I think a better analogy would be if you suddenly, even though it is not causing you trouble, denied access to someone picking up food that you didn't contract with, with the full knowledge that the someone would be in big trouble with their customers.
Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]
#219Earlier quoted context omitted.
Not allowing the CFAA to be (ab)used to attempt to make scraping illegal makes sense. However, how is it reasonable to force a web site to serve its contents to a third-party company, without being allowed to make a decision whether to serve it or not? Serving the web site costs money, and the scraper surely isn't going to generate ad income...
> However, how is it reasonable to force a web site to serve its contents to a third-party company, without being allowed to make a decision whether to serve it or not? Your statement makes absolutely no sense. That's not how internet works. If you serve something publicly you don't get to cherry pick who sees it. Not only it makes no sense technically it's also a huge anti-competitive case.
Technically, of course you can identify IP ranges owned by certain entities and restrict their access. That’s trivial, so what do you mean when you say the internet doesn’t work like that?
Legally, there’s plenty of region locked content for copyright and censorship reasons. A distributor might region lock because they don’t have distribution rights in particular regions. Are you saying distributors can’t publish free content at all because they can’t choose who sees it but would be breaking copyright law to publish to everyone? Or a site might region lock because certain content is censored in particular countries. Can you not publish anti-regime articles because a totalitarian country is on the Internet?
The entire world isn’t and shouldn’t be held hostage to the most restrictive laws that exist in the world. And the answer isn’t blocking on the requesting end because that’s technically much harder and blocks much, much more content. So what am I missing?
Edit: Forgot to include the other end of the spectrum. If I, as an individual, host my own site on my own hardware with my own connection that I pay the bandwidth for, can I deny a suspected not network?
Re: 9th Circuit holds that scraping a public website does not violate the CFAA [pdf]
#220Earlier quoted context omitted.
For sure, but they shouldn't be hypocritical about it. If they don't consider themselves content parasites, they shouldn't consider people scraping their site to be content parasites, either. (Some sites really are just parasites, though.)
There's nothing hypocritical about it. Googlebot respects robots.txt configured on pages it scrapes. Google in turn expects that their own robots.txt will be respected. What's the issue? https://www.google.com/robots.txt
If you want to talk about this in terms of robots.txt, Google is thriving on the fact that other companies don't block their content in robots.txt, but at the same time Google blocks all of its content in its robots.txt.