Live data from Hacker News

The Scraping Problem and Ethics

blog.osvdb.org

121–130 of 130 posts

Re: The Scraping Problem and Ethics

#121

This is one of the more interesting policy questions on the web. Our search engine crawls a lot of blogs and what not on the web, criminals who want to find unpatched wordpress sites try to acrape our crawl by sending automated (scripted) queries to find them. We have developed a number of defenses over the years and pretty regularly ban them[1]. Here is the weird part though, if they hired 300 people on mechanical t…

I think there is a great meta-question in here, about business models for digital data and software. Here you have a great case study, about an organization that tried to do a volunteer model, and it didn't work. Then they pivoted to a commercial model, but fundamentally they still believe in a free tier. But they have to cripple that free tier pretty thoroughly, and even still people abuse it. I have a product I'm w…

Perhaps it won't come as a surprise but I've been toying with this question for a couple of decades now. Specifically what are the economics of information? In the 'goods' economy there are some interesting mechanisms that inform the question of value, these include but are not limited to, the cost to acquire materials, develop expertise in manufacturing, and managing the supply lines between raw material to finished good. Accountants will talk about the "Cost of goods sold" as a grouping function for these costs. In the 'information' economy the manufacturing part it pretty trivial, you just replicate copies, but the assembling part can be quite difficult. This leads to an interesting inversion where it can cost a lot to assemble something and nothing to 'manufacture' it. And that doesn't even begin to touch on what it is about information that makes it valuable in the first place.

What is the difference in value between a CD with the latest release of Ubuntu burned on it, and the download? download and a bootable Flash drive?

There is a great experiment you can run which goes like this; At one end of an athletic field, place a chess board with a queen on it on one of the squares. At the other end of the field have a table where people can get a quest. Offer to pay a person $5 if they will walk to the end of the field, note where the queen is, and come back and tell the quest giver. At the mid point of the field set up an information seller. They offer to sell you the location of the queen for anywhere between 20 and 80% of the reward price.

This simple experiment lets you see all sort of mechanisms in play that control information value. On the one hand you can see the range of values people apply to their own time (acquisition cost), their willingness to retain value (do they then go back mid-queue at the sign up table and start offering to sell the information for some fraction of the price to anyone?) At what threshold to people start trying to break the rules (a notion that is similar to price inelasticity but has a component like the 'black market demand').

Interesting questions to be sure.

Re: The Scraping Problem and Ethics

#122
post #16

Earlier quoted context omitted.

So many times I say this to people. Sometimes the reason people aren't buying is because they can't find the price, and aren't willing to pay the mental price of talking to a sales person to find out. I understand the sales psychology in making sure you enter into a proper discussion with people to make sure their needs are right, and extracting the maximum consumer surplus from them. But there is a non-trivial prici…

I get this line of thinking if your you can sell your offering for $100 bucks a month. But what if your minimum offering is $2500/mo? Correct me if I'm wrong, but I've yet to see a company - that only sales through enterprise - put up a sign saying where their plans start at $2,500. In most cases I don't believe these services are forcing customers to talk on the phone because they think they will convert them. Chanc…

Yes, that's what I am saying. There is a pricing point at which having a customized sales process absolutely makes sense.

What I am saying along with that is that - by doing this - you're ignoring a lot of the market at (or even just below) your price point. That might be OK - as long as it is a conscious decision to do so, including acknowledging that your market is completely above that point.

In the case being discussed here, it would seem that there is an interest below this price point.

Re: The Scraping Problem and Ethics

#123
post #47

Earlier quoted context omitted.

Tell that to the 95% of the population on this site who torrent TV shows, movies and music.

Even ignoring the difference between taking something for personal use and taking something for commercial use (as that is a distinction that does not affect the legality of the matter) "because other people do it" is not a valid defence.

It's not a defence, I think you can see from my tone I don't support stealing media just because you don't like how it is distributed.

I'm saying rather this argument won't find much favour here unless everybody is a hypocrite.

Re: The Scraping Problem and Ethics

#124
post #96

Earlier quoted context omitted.

Pot calling the kettle black. Crawling others' websites then selling ads. Scraping others' websites then selling ads. Selling access to user-generated content. Amusing to watch these folks argue about ethics. Who owns the copyrights in this data? Surely not the one who is demanding that you pay a license fee. These "services" are middlemen, plain and simple. This might be why McAfee was wondering about how much manua…

If they're a middleman, and it's such an easy "service", why didn't McAfee just bypass them and get the data from the original sources, rather than do the wrong thing?

That's a valid question and one I have considered myself.

So what drives the folks at McAfee to do this?

Maybe it is the same thinking that drives programnmers to not want to write code.

"Don't reinvent the wheel."

"Code reuse."

"Use a shared library or a scripting language with batteries included."

Why crawl the web when Google has already done it for you?

And so on.

Personally, I do not have trouble understanding why McAfee would do this.

What I have trouble understanding is why OSVDB would think they could crawl some public data and then charge a fee to access it.

It is the "sale" of "free" information that puzzles me.

By all means go ahead and try, you may well recoup your outlays for compiling the free data and even make a profit.

But should we really be surprised when someone does not want to pay?

Re: The Scraping Problem and Ethics

#125
post #73

Earlier quoted context omitted.

I'm assuming that the vast majority of the time, it's a case of an individual or small group that wants the information, not a real corporation. Usually they have no intention of paying for anything or playing by the rules. I have absolutely no explanation regarding McAfee though, considering the billions McAfee and its parent company, Intel, makes in revenue yearly.

Could this be a company culture thing at McAfee ?

I sure hope not.

Re: The Scraping Problem and Ethics

#126

This is one of the more interesting policy questions on the web. Our search engine crawls a lot of blogs and what not on the web, criminals who want to find unpatched wordpress sites try to acrape our crawl by sending automated (scripted) queries to find them. We have developed a number of defenses over the years and pretty regularly ban them[1]. Here is the weird part though, if they hired 300 people on mechanical t…

Looks like 80legs that you mention rebranded themselves as "Datafiniti" at some point recently. http://blog.datafiniti.net/?p=230

which in Italian literally means 'end of Data'... kind of ironic isn't it?

Re: The Scraping Problem and Ethics

#127
post #68

Earlier quoted context omitted.

>More players in the market means more competition and more competition means a better service I'm not sure if this is a joke. So, I'll refrain from replying.

Could you explain your point of view? I'd like to understand both sides here and I think I understand why more competition would be good, but now why it wouldn't be.

>I think I understand why more competition would be good, but now why it wouldn't be.

That is not what I said at all. I don't believe that more competition necessarily leads to better product/service. In fact I believe that in most cases it does not. In capitalist economies, companies try their hardest to avoid competing. Competition forces companies to reduce costs, and not necessarily increase quality. The quality of a product is not some number that people can read and go "oh yeah this product is better". Marketing people try hard to invent such pointless numbers (e.g. Megapixels in cameras.. horsepower in cars , etc etc). Also many CEOs don't have the first clue on how the product is actually made, much less increase its quality. They rely on these same 'marketing numbers' that their underlings serve them with. So, they too can go "oh yeah this number is increasing so our product is getting better".

It seems like a lot of people are brainwashed into believing this free market utopia where things just automatically get better because everyone is competing and the customer is this genius who can figure out which company is delivering a better product.

Its sort of like thinking "Well if I'm nice to everyone, everyone will be nice to me.". And then you realize that the real world is a dark place filled with assholes, where slavery is still rampant and many of the goods and services we consume are dependent on the exploitation of natural resources or other fellow humans.

Sorry if reading all that bummed you out. I'm really a quite a cheerful person :P

Re: The Scraping Problem and Ethics

#129
post #112

Earlier quoted context omitted.

but how is it stealing when one does not lose inventory? If one person scrapes a page, did you lose the source code? Does it not become available for the next visitor? What possible loss do you incur that is directly tied to your data? When you make data public with the intent of being readily accessible by the public, how can you claim theft when you are achieving what you set out to do? Does the accelerated rate of…

"but how is it stealing when one does not lose inventory?" Scraping is not necessarily a no-victim situation. Even today after this stuff has gotten cheaper, you're costing them bandwidth fees, and likely increasing their server storage and CPU fees if it's on a metered hosting service, which is quite likely nowadays. If you degrade their site's functionality, you may chase away paying customers. We need not hypothes…

> Scraping is not necessarily a no-victim situation. Even today after this stuff has gotten cheaper, you're costing them bandwidth fees, and likely increasing their server storage and CPU fees if it's on a metered hosting service, which is quite likely nowadays. If you degrade their site's functionality, you may chase away paying customers.

While technically correct, you are conflating the issues, because in none of the cases (that I've seen mentioned so far in this thread) the problem is with bandwidth/storage/CPU costs of retrieval to any significant extent.

Instead, it appears that almost all of the costs are incurred before retrieval: curating, sorting, etc.

I'm not arguing that it's okay, but it's just as much not stealing / thievery as downloading movies or music isn't.

Re: The Scraping Problem and Ethics

#130
post #43
post #42

Earlier quoted context omitted.

>You might argue the difference is that the information that Aaron was after was already paid for with public money. This is naive. Just because research is backed by public money doesn't mean the publications are automatically free to the public. If your argument was valid, you could use it to demand access to the emails of every FBI employee. Just because something is funded by the public doesn't automatically make…

http://en.wikipedia.org/wiki/Freedom_of_Information_Act_%28U... http://en.wikipedia.org/wiki/Government_in_the_Sunshine_Act http://en.wikipedia.org/wiki/Brown_Act

Have to file a foia request that could take years and be denied isn't exactly open.
Post reply on HN