Live data from Hacker News

The Scraping Problem and Ethics

blog.osvdb.org

81–90 of 130 posts

Re: The Scraping Problem and Ethics

#81

This is one of the more interesting policy questions on the web. Our search engine crawls a lot of blogs and what not on the web, criminals who want to find unpatched wordpress sites try to acrape our crawl by sending automated (scripted) queries to find them. We have developed a number of defenses over the years and pretty regularly ban them[1]. Here is the weird part though, if they hired 300 people on mechanical t…

I think there is a great meta-question in here, about business models for digital data and software. Here you have a great case study, about an organization that tried to do a volunteer model, and it didn't work. Then they pivoted to a commercial model, but fundamentally they still believe in a free tier. But they have to cripple that free tier pretty thoroughly, and even still people abuse it. I have a product I'm w…

perfect price discrimination is hard.

Re: The Scraping Problem and Ethics

#82
post #69
post #13

Earlier quoted context omitted.

like much security, it's not about making it impossible, it's about making it a lot less convenient/a bit harder. At one point the effort to circumvent would cost more in man-hours than just buying the product.

You can get scrapping libraries fairly easily. In my more shady past I developed a library like that and shared it, HTTP client library with automatic proxy rotation and rate limiting friendly. When I used it (which was almost a decade ago) never ran into problems, plug a list of 10,000 proxy and scrap away. Not condoning that, which is a bit hypocrite of me, at the time I was mostly doing what I was told and I thoug…

I know that paying for things was annoying when I didn't have money. Now that I have cash, I'm willing to pay (reasonable amounts of ) money for digital things.

I still hesitate when it comes to thousand dollar licenses when it's for my personal use , though.

Re: The Scraping Problem and Ethics

#83

Earlier quoted context omitted.

You need to be careful sending bogus data: in some jurisdictions this could be argued to be deliberate targeted commercial sabotage. You would no doubt eventually win any resulting legal argument, assuming you could afford to carry the argument on to that conclusion . Sending no data, or limited data, would be safe though. A better method would be to set "default" pricing (something high but not ridiculous, that coul…

>" You need to be careful sending bogus data: in some jurisdictions this could be argued to be deliberate targeted commercial sabotage. " // That sounds pretty ludicrous, do you have anything to back it up - caselaw, settlement report? It would be analagous to serving a fake image to combat hotlinking; or a fake page to combat framing.

Something that is obviously fake or otherwise different (like the image or frame break-out examples) would also be fine.

But leading someone to believe they have correct data when what they have is potentially embarrassing when used could be something they'd take objection to. Even if not there are two other points of risk: your reputation if something goes wrong and you accidentally give bad data to your paying clients and your reputation if someone, paying or otherwise, shows off the bad data as "the sort crap these people try to sell".

I don't have any specific references, but it is something I would be careful of as there have certainly been similarly ludicrous (IMO) cases on unrelated matters in the past (yes the right side would win, assuming they can afford to).

I may be being too cynical here, then again maybe not...

I did misread the grand-parent post though and this isn't what he was talking about. He was suggesting the bogus data was the message, and I read it as handing out fake data for the scraper along with the message to be seen should a human be looking.

Re: The Scraping Problem and Ethics

#84

Earlier quoted context omitted.

You need to be careful sending bogus data: in some jurisdictions this could be argued to be deliberate targeted commercial sabotage. You would no doubt eventually win any resulting legal argument, assuming you could afford to carry the argument on to that conclusion . Sending no data, or limited data, would be safe though. A better method would be to set "default" pricing (something high but not ridiculous, that coul…

You don't have to return bogus data, you can return an HTTP error code: 402 payment required.

Exactly.

The problem I see is with giving out bad data while leading people to believe they have obtained useful information (that they then embarrass themselves by using/re-distributing).

I did misread the grand-parent post though: his was suggesting the bogus data was the message, and I read it as handing out fake data for the scraper along with the message to be seen should a human be looking.

Re: The Scraping Problem and Ethics

#86

Earlier quoted context omitted.

You don't have to return bogus data, you can return an HTTP error code: 402 payment required.

Exactly. The problem I see is with giving out bad data while leading people to believe they have obtained useful information (that they then embarrass themselves by using/re-distributing). I did misread the grand-parent post though: his was suggesting the bogus data was the message, and I read it as handing out fake data for the scraper along with the message to be seen should a human be looking.

Yeah I can see why fake data could give you a harder time in terms of lawsuits. Similar to people who fight hotlinking of their content by replacing images with something offensive.

Re: The Scraping Problem and Ethics

#87

This is one of the more interesting policy questions on the web. Our search engine crawls a lot of blogs and what not on the web, criminals who want to find unpatched wordpress sites try to acrape our crawl by sending automated (scripted) queries to find them. We have developed a number of defenses over the years and pretty regularly ban them[1]. Here is the weird part though, if they hired 300 people on mechanical t…

> Look at folks like 80legs or other 'distributed' scrapers. They exist almost solely to subvert these service terms.

Not sure I understand, is there a whole ecosystem of web businesses that feed off for free of your search engine, or you meant from several legit search engines or ... ?

Re: The Scraping Problem and Ethics

#88

This is one of the more interesting policy questions on the web. Our search engine crawls a lot of blogs and what not on the web, criminals who want to find unpatched wordpress sites try to acrape our crawl by sending automated (scripted) queries to find them. We have developed a number of defenses over the years and pretty regularly ban them[1]. Here is the weird part though, if they hired 300 people on mechanical t…

osvdb.org has moved from 72.233.69.6 to 76.74.254.123

Re: The Scraping Problem and Ethics

#89

This is one of the more interesting policy questions on the web. Our search engine crawls a lot of blogs and what not on the web, criminals who want to find unpatched wordpress sites try to acrape our crawl by sending automated (scripted) queries to find them. We have developed a number of defenses over the years and pretty regularly ban them[1]. Here is the weird part though, if they hired 300 people on mechanical t…

I think there is a great meta-question in here, about business models for digital data and software. Here you have a great case study, about an organization that tried to do a volunteer model, and it didn't work. Then they pivoted to a commercial model, but fundamentally they still believe in a free tier. But they have to cripple that free tier pretty thoroughly, and even still people abuse it. I have a product I'm w…

They made the process of paying for the software a laborious pain in the ass. They are desperately trying to extract money from those who can pay, which sadly drives away those who can pay but don't want an involved process.

I have worked at lots of companies where I had a monthly budget of 10k+ that I could spend on whatever I wanted, but if I wanted any sort of complex deal (can't just put on CC with a line item) -- had to bring in legal and other groups -- instantly killed any interest.

"Licensing is based on the data needed (e.g. all of it vs subset), how it is used (e.g. internal only, external, product integration), etc."

What a goddamn horror show. I simply want a product, I want to pay for it, and I want to use it. Turning on Dropbox for Business was a decision made in about 5 minutes... "You all like it, already using it, awesome! I will get team setup." -- 5 minute later I had given Dropbox $3800.

I really think they are getting in their own way for no benefit. They have created a very high barrier to EVEN HAVING A DISCUSSION about buying the product. So, if I don't know exactly how will use it -- I can't purchase it. Stupidity.

Re: The Scraping Problem and Ethics

#90
post #77

This is one of the more interesting policy questions on the web. Our search engine crawls a lot of blogs and what not on the web, criminals who want to find unpatched wordpress sites try to acrape our crawl by sending automated (scripted) queries to find them. We have developed a number of defenses over the years and pretty regularly ban them[1]. Here is the weird part though, if they hired 300 people on mechanical t…

Since you work at Blekko, there are more points to discuss that were not addressed in the article. For example: i) Are search engines web scrapers? ii) Should search engines pay the scraped sites if they are charging to access their indexed data? probably some of the scraped sites has a specific license forbidding the search engine to sell their information in any way. iii) Regarding Internet policies, is it fair/unf…

> ii) Should search engines pay the scraped sites if they are charging to access their indexed data? probably some of the scraped sites has a specific license forbidding the search engine to sell their information in any way.

Any reputable search engine will respect robots.txt.

Post reply on HN