Web scraping isn't a crime - the simple act of downloading the data should not be a problem here. (The reuse of the data might be, depending, but we don't have that information right now.) This doesn't even rise to the level of what Weev did.
He's just making the point that it's unethical, which it is. Even if it boils down to simple bandwidth theft.
The Scraping Problem and Ethics
61–70 of 130 posts
Re: The Scraping Problem and Ethics
#62Earlier quoted context omitted.
It's not really reasonable to say "I don't like the way you market your goods... so you really shouldn't be concerned with people stealing them."
Tell that to the 95% of the population on this site who torrent TV shows, movies and music.
Re: The Scraping Problem and Ethics
#63I'm more and more concerned that the legal and cultural environment for web scraping would make it hard for a company like Google or Yahoo to be founded today. The internet isn't about "don't take my stuff", it's about spreading that stuff around. I'm confused by people who want to make their data public, but want to control exactly how people access it.
Is that a bad thing?
>The internet isn't about "don't take my stuff", it's about spreading that stuff around.
Try asking Google if they want to "share" their database of crawled data.
>I'm confused by people who want to make their data public, but want to control exactly how people access it.
Me too.
Re: The Scraping Problem and Ethics
#64If I were in OSVDB's shoes, I would call these people out in an email and ask them to pay a licencing fee. McAfee have always had a shady past, even when they shook themselves clean of John, they have a history of scam-like behaviour to make a quick-buck. You should be rate-limiting how many requests free API users can make, like; Twitter, Facebook and every other Internet provider does via their API. Make it harder…
Re: The Scraping Problem and Ethics
#65Earlier quoted context omitted.
[deleted]
Actually, Terms of Service violations fall under the Computer Fraud and Abuse Act, since ToS agreements can lay out under which circumstances that authorization for access to computer systems is given. That sort of obscene generality is the reason for proposals such as Aaron's Law, but to my knowledge there are no such protections today.
Re: The Scraping Problem and Ethics
#66I'm more and more concerned that the legal and cultural environment for web scraping would make it hard for a company like Google or Yahoo to be founded today. The internet isn't about "don't take my stuff", it's about spreading that stuff around. I'm confused by people who want to make their data public, but want to control exactly how people access it.
>make it hard for a company like Google or Yahoo to be founded today. Is that a bad thing? >The internet isn't about "don't take my stuff", it's about spreading that stuff around. Try asking Google if they want to "share" their database of crawled data. >I'm confused by people who want to make their data public, but want to control exactly how people access it. Me too.
Re: The Scraping Problem and Ethics
#67If I were in OSVDB's shoes, I would call these people out in an email and ask them to pay a licencing fee. McAfee have always had a shady past, even when they shook themselves clean of John, they have a history of scam-like behaviour to make a quick-buck. You should be rate-limiting how many requests free API users can make, like; Twitter, Facebook and every other Internet provider does via their API. Make it harder…
If I were in their shows I would track their IPs and send them bogus data along the lines of "Please pay for a commercial license."
A better method would be to set "default" pricing (something high but not ridiculous, that could easily be negotiated downwards if they contact you) and make access beyond a few requests a click-through (or better: have them respond to an email before progressing further) where they agree to that pricing if they are using the information commercially.
Re: The Scraping Problem and Ethics
#68Earlier quoted context omitted.
>make it hard for a company like Google or Yahoo to be founded today. Is that a bad thing? >The internet isn't about "don't take my stuff", it's about spreading that stuff around. Try asking Google if they want to "share" their database of crawled data. >I'm confused by people who want to make their data public, but want to control exactly how people access it. Me too.
>Is that a bad thing? Yes, no question about it. More players in the market means more competition and more competition means a better service.
I'm not sure if this is a joke. So, I'll refrain from replying.
Re: The Scraping Problem and Ethics
#69Earlier quoted context omitted.
On the other hand, large (even non-profit) organisations are precisely the ones who have the resources to scrape steathily and widely, as would a loosely-organised community of users... it's not hard to come up with algorithms to respect the rate limits, balancing the load across multiple IPs and accounts, and producing access patterns that don't look any different from the rest of the site traffic.
like much security, it's not about making it impossible, it's about making it a lot less convenient/a bit harder. At one point the effort to circumvent would cost more in man-hours than just buying the product.
When I used it (which was almost a decade ago) never ran into problems, plug a list of 10,000 proxy and scrap away.
Not condoning that, which is a bit hypocrite of me, at the time I was mostly doing what I was told and I thought I was clever. Now that I'm in a position to have a positive impact, I do buy data and pay appropriate licence fees on all software/data purchase, which still baffles some of my programmers who constantly ask "why not crack it?", "you know I found a .zip on Google with the data, why buy it?", and so forth.
I don't know what in programmer culture makes it so hard for us to pay for something, some people put some effort behind that software / data collection, and it's only fair to pay them.
Re: The Scraping Problem and Ethics
#70Earlier quoted context omitted.
Maybe they want Google to be able to crawl their database (which it has clearly done, as you'll see if you do a search.) That also raises some questions...
Not completely by the specification, but I think this one works as expected. user-agent: * disallow: / user-agent: Googlebot allow: /