http://blog.wired.com/wiredscience/2008/01/google-to-provi.h...
I am not sure if this is true or what the status is.
Open Data Commons has an interesting license for open datasets
21–30 of 56 posts
http://blog.wired.com/wiredscience/2008/01/google-to-provi.h...
I am not sure if this is true or what the status is.
Open Data Commons has an interesting license for open datasets
It's so painfully obvious that this is a good idea, and yet people are skeptical of the Semantic Web, which has been promoting this idea for almost a decade (and doesn't require centralization). Why?
http://www.well.com/~doctorow/metacrap.htm
http://www.shirky.com/writings/semantic_syllogism.html
http://blahsploitation.blogspot.com/2005/09/i-always-figured...
It's like freebase all right, but the article has something right about adoption. He points out some big company would have to donate a great starting dataset to drive adoption. I think this is one problem with Freebase. Another problem I see is the structuring of the data, it is a hurdle to sharing. Finally the last problem I see is Data is currently very much seen as a competitive advantage. When Google introduce f…
It's like freebase all right, but the article has something right about adoption. He points out some big company would have to donate a great starting dataset to drive adoption. http://news.ycombinator.com/item?id=157966
One of the biggest competitive advantages companies have is data. Like the article and previous comment already said, adoption is the hardest barrier. Unless someone can provide a compelling reason for companies/scientists/etc to give data or data access, there really isn't much else to discuss.
I think sites like mashery and dapper.net are going in the right direction, by providing good licensing rights and monetization controls that can incent large companies with reliable datasets to participate.
It's so painfully obvious that this is a good idea, and yet people are skeptical of the Semantic Web, which has been promoting this idea for almost a decade (and doesn't require centralization). Why?
Here's three anti-SemWeb articles that try to answer your question: http://www.well.com/~doctorow/metacrap.htm http://www.shirky.com/writings/semantic_syllogism.html http://blahsploitation.blogspot.com/2005/09/i-always-figured...
The Shirky article is a well-known strawman.
Thanks for the pointer to the last one, I'll read it when I get a chance.
http://en.wikipedia.org/wiki/The_Thomson_Corporation
Could they really end up being the Encarta Encyclopedia of data?
Two years ago I wrote http://formula1db.com , to teach myself sql. To get the data, I had to screen scrape formula1.com . Once I had the data, learning SQL became a joy. I haven't done much with the data since I built the site, but I am considering open sourcing it. What I don't understand is, why don't sport leagues open source their data. What do they lose? Its a good think that people are so excited about your spo…
http://baseball1.com/content/view/57/82/ http://retrosheet.org/
Current major and minor league data is available as well, though MLB will crack down on anyone who is trying to make money off of derivative products. Here's where you'll find it, as XML:
http://gdx.mlb.com/components/game/
One can do pretty cool stuff with all of it, and many people have, despite the fact that we can't make money off of it:
http://minorleaguesplits.com/ (my site)
http://baseball.bornbybits.com/2008/pitchers.html (analysis based on detailed pitch speed / break information that MLB started collecting last year.)
Joel had an article on "commoditizing your complements": http://www.joelonsoftware.com/articles/StrategyLetterV.html . Of course we want to commoditize data, to raise the value of hackers (raise the expected return from hacking). On the other side, data companies want to commoditize hackers, to raise the value of data... a process we are sure to resent.
It's so painfully obvious that this is a good idea, and yet people are skeptical of the Semantic Web, which has been promoting this idea for almost a decade (and doesn't require centralization). Why?
But, this is also not only a technology solution. People have to make a choice themselves to filter and promote good information.