Live data from Hacker News

UK chemist on Elsevier's ban on textmining

blogs.ch.cam.ac.uk

1–10 of 30 posts

Re: UK chemist on Elsevier's ban on textmining

#2
Until the academia realizes that computers can read too, it's important for companies that DO have access to these papers (like google scholar[1]) to create 3rd party APIs so that we can at least have better search tools.

1. http://code.google.com/p/google-ajax-apis/issues/detail?id=1...

Re: UK chemist on Elsevier's ban on textmining

#3
I'm not sure how a tool to read some text and display a diagram of the chemical reaction described falls under the "law" of the quoted passage. Copyright law certainly allows for this, so all the journals can do is say, "you don't get to buy our feed anymore if you run your program on our articles," but surely this is a game of chicken because no company wants to lose tens of thousands of dollars a year for no reason.

I would implement this as a browser plugin that uploads the content to a server (like Google Translate), and let the journals deal with each rogue user individually.

Telling people what software they can use to read text doesn't scale.

Re: UK chemist on Elsevier's ban on textmining

#5
post #3

I'm not sure how a tool to read some text and display a diagram of the chemical reaction described falls under the "law" of the quoted passage. Copyright law certainly allows for this, so all the journals can do is say, "you don't get to buy our feed anymore if you run your program on our articles," but surely this is a game of chicken because no company wants to lose tens of thousands of dollars a year for no reason…

What would happen if you actually did try to textmine it. surely this is a game of chicken because no company wants to lose tens of thousands of dollars a year for no reason.

Elsevier is willing to play this game of chicken; they are convinced that their customers (i.e., research universities) cannot do without their product. Hence the OP reporting that twice, they instantly and without warning cut off all access to their product to the entire university, because of detected scraping.

Telling people what software they can use to read text doesn't scale.

Unfortunately, I think it does. If individual users, even a large number of them, use a plug-in, they might get away with it-- but this would result in an incomplete data set. To systematically textmine the corpus (which is the task at hand) requires some kind of systematic access to the data, and this is where Elsevier steps in and shuts it down.

Re: UK chemist on Elsevier's ban on textmining

#6
post #3

I'm not sure how a tool to read some text and display a diagram of the chemical reaction described falls under the "law" of the quoted passage. Copyright law certainly allows for this, so all the journals can do is say, "you don't get to buy our feed anymore if you run your program on our articles," but surely this is a game of chicken because no company wants to lose tens of thousands of dollars a year for no reason…

What would happen if you actually did try to textmine it. surely this is a game of chicken because no company wants to lose tens of thousands of dollars a year for no reason. Elsevier is willing to play this game of chicken; they are convinced that their customers (i.e., research universities) cannot do without their product. Hence the OP reporting that twice, they instantly and without warning cut off all access to…

Possibility number two is to break in to a router closet at MIT to do the scraping :)

Re: UK chemist on Elsevier's ban on textmining

#7
post #6

Earlier quoted context omitted.

What would happen if you actually did try to textmine it. surely this is a game of chicken because no company wants to lose tens of thousands of dollars a year for no reason. Elsevier is willing to play this game of chicken; they are convinced that their customers (i.e., research universities) cannot do without their product. Hence the OP reporting that twice, they instantly and without warning cut off all access to…

Possibility number two is to break in to a router closet at MIT to do the scraping :)

What actually happen with that guy? I remember all the mainstream media stories what kind of a bad ass hacker he was, but somewhat never caught up with the "happy end".

Re: UK chemist on Elsevier's ban on textmining

#8
post #7
post #6

Earlier quoted context omitted.

Possibility number two is to break in to a router closet at MIT to do the scraping :)

What actually happen with that guy? I remember all the mainstream media stories what kind of a bad ass hacker he was, but somewhat never caught up with the "happy end".

http://en.wikipedia.org/wiki/Aaron_Swartz#JSTOR

Re: UK chemist on Elsevier's ban on textmining

#9
post #7

Earlier quoted context omitted.

What actually happen with that guy? I remember all the mainstream media stories what kind of a bad ass hacker he was, but somewhat never caught up with the "happy end".

http://en.wikipedia.org/wiki/Aaron_Swartz#JSTOR

Not really a bad ending. The feds have no case against him, MIT and JSTOR aren't going to sue him, and JSTOR decided that releasing their archive was a good idea. Hopefully the FBI will drop their case.

Re: UK chemist on Elsevier's ban on textmining

#10

It's not the semantic interpretation of text that they're banning, but the scraping of text, which deals with copying what they'd prefer to sell you. Still an abhorrence, but let's get our facts straight.

Actually, it sounds like they are trying to sell their interpretation of a recipe because they recently acquired a company which extracts recipes. They've banned their customers from running the same tool which they probably use internally in order to protect their slice of the market.
Post reply on HN