Earlier quoted context omitted.
It would be ok if it wasn't "You can't scrap my site. Unless of course you're Google" this double standard drives me mad.
As the owner of a large website, I don't care what you think. I block by default and whitelist when I decide it's in my interest. If you don't think this is reasonable, chances are you've never run a large website, or analyzed the logs of a large website. You'd be astonished how much robotic activity you'll receive. If left unchecked it can easily swamp legitimate traffic. Unless you have a way for me to automaticall…
Web Scraping in 2016
301–310 of 402 posts
Re: Web Scraping in 2016
#302They consider their data to be theirs, even though they published it on the internet. They consider your data (your personal integrity) to be theirs as well, because how can you assume personal integrity when you are surfing the internet?
I have high hopes that the judicial system some time not too far from now will realize that since the law should be a reflection of the current moral standings it will always be behind, trying to catch up with us and that those who break the law while not breaking the current moral standings are still "good citizens" unworthy of prison or fines.
I guess Google won this iteration of the internet because of the double-standars site owners stand by, to allow Google to scrape anything while hindering any competitors from doing the same. There will only be a true competitor to Google when we in the next iteration of the internet realize that searching vast amounts of data (the internet) is a solved problem, that anyone can do as good a job as Google, and move on to the next quirk, around wich there will be competition, and in the end that quirk will be solved, we'll have a winner, signaling that is it time to move on to the next iteration.
Re: Web Scraping in 2016
#303Earlier quoted context omitted.
A majority of the websites that blekko, a google competitor, contacted to ask for robots.txt access ignored us.
I agree, it's a difficult conundrum. It sucks.
Re: Web Scraping in 2016
#304Earlier quoted context omitted.
As the owner of a large website, I don't care what you think. I block by default and whitelist when I decide it's in my interest. If you don't think this is reasonable, chances are you've never run a large website, or analyzed the logs of a large website. You'd be astonished how much robotic activity you'll receive. If left unchecked it can easily swamp legitimate traffic. Unless you have a way for me to automaticall…
As the user of large websites I don't care. I'm not going to read the TOS and I will continue to scrape what I like since it makes my life more convenient. Like OP when blocked I'll just drive my scraping through a web browser which is the same as I've done for years on various sites that never provided APIs.
Re: Web Scraping in 2016
#305Anyways I added your stuff here along with other data mining resource:
https://github.com/kevindeasis/awesome-fullstack#web-scrapin...
Re: Web Scraping in 2016
#306Earlier quoted context omitted.
And the original point of my comment was that doing this is extremely rude and not appropriate, not that it couldn't be done or that others weren't doing it. Feel free to send any request to any server you want, it is certainly up to them to decide whether or not to serve it, but that doesnt absolve you of guilt from scraping someone's site when they explicitly ask you not to.
Please don't conflate "extremely rude", "not appropriate", and "guilt". Two of these are subjective opinions about what constitutes good citizenship. The last one is a legal determination that has the potential to deprive an individual of both his money and liberty. We're discussing whether these behaviors should be legal , not whether they are necessarily polite.
You are posting in a comment thread underneath my reply about rudeness and impoliteness, ironically being somewhat rude telling me off about what not to conflate when it was never what I said.
Re: Web Scraping in 2016
#307Earlier quoted context omitted.
It's different because you don't contact another party's server to do it. The CFAA makes it illegal to "exceed authorized access" to networked computers. "Authorized access" is whatever the server's owner says it is. That's why the copyright status of factual accumulations isn't a protection for internet scraping. If Feist v. Rural occurred now and Rural, like most companies, kept their information in a database onli…
You've made a large number of authoritative-sounding comments on this story... and this one, like many of the others, is a guess.
Re: Web Scraping in 2016
#308Earlier quoted context omitted.
It would be ok if it wasn't "You can't scrap my site. Unless of course you're Google" this double standard drives me mad.
Worse is that Google tries to stop scraping. It's like they don't want anyone to see past the first page of results. They could scrape your website and then they prevent you form scraping your own data back. The whole process is silly; it reflects the duct tape and chicken wire nature of the www. No one should have to "scrape" or "crawl". Data should be put into a open universal format (no tags) and submitted when ne…
Re: Web Scraping in 2016
#309Keep in mind that companies have sued for scraping not through the API, for example LinkedIn, which explicitly prevents scraping via the ToS: http://www.informationweek.com/software/social/linkedin-sues... OKCupid did a DMCA takedown for researchers releasing scraped data: https://www.engadget.com/2016/05/17/publicly-released-okcupi... Since both of these incidents, I now only scrape if it's a) through the API follow…
On the ethics side, I don't scrape large amounts of data - eg. giving clients lead gen (x leads for y dollars) - in fact, I have never done a scraping job and don't intend to do those jobs for profit. For me it's purely for personal use and my little side projects. I don't even like the word scraping because it comes loaded with so many negative connotations (which sparked this whole comment thread) - and for a good…
You are exactly right. But although a site can deny you access for any arbitrary reason (it's their website, after all) obviously government think they are the ones to enforce this crap.
What if the ToS say you can only access a site while jumping hoops? Only read the ToS after a while and wasn't hooping? Well too bad, now you are being sued for reading the main page _and_ the ToS page without jumping around.
This comment Terms of Service: If you read any of this text you owe lerpa $1.000.000 to be paid up until 09/01/2016.
Re: Web Scraping in 2016
#310Corporations will abuse your personal integrity whenever they get a chance, while abiding the law. Corporations will cry like babies when their publicly available data (their livelyhood) gets scraped. They will take you to court. They consider their data to be theirs, even though they published it on the internet. They consider your data (your personal integrity) to be theirs as well, because how can you assume perso…