Live data from Hacker News

We need a Wikipedia for data

bret.appspot.com

21–30 of 56 posts

Re: We need a Wikipedia for data

#22
post #15

It's so painfully obvious that this is a good idea, and yet people are skeptical of the Semantic Web, which has been promoting this idea for almost a decade (and doesn't require centralization). Why?

Here's three anti-SemWeb articles that try to answer your question:

http://www.well.com/~doctorow/metacrap.htm

http://www.shirky.com/writings/semantic_syllogism.html

http://blahsploitation.blogspot.com/2005/09/i-always-figured...

Re: We need a Wikipedia for data

#23
post #13
post #6

It's like freebase all right, but the article has something right about adoption. He points out some big company would have to donate a great starting dataset to drive adoption. I think this is one problem with Freebase. Another problem I see is the structuring of the data, it is a hurdle to sharing. Finally the last problem I see is Data is currently very much seen as a competitive advantage. When Google introduce f…

It's like freebase all right, but the article has something right about adoption. He points out some big company would have to donate a great starting dataset to drive adoption. http://news.ycombinator.com/item?id=157966

This improves the (easy) technical hurdles, but how about (very difficult) social, economic, and bureaucratic ones?

One of the biggest competitive advantages companies have is data. Like the article and previous comment already said, adoption is the hardest barrier. Unless someone can provide a compelling reason for companies/scientists/etc to give data or data access, there really isn't much else to discuss.

I think sites like mashery and dapper.net are going in the right direction, by providing good licensing rights and monetization controls that can incent large companies with reliable datasets to participate.

Re: We need a Wikipedia for data

#24
post #15

It's so painfully obvious that this is a good idea, and yet people are skeptical of the Semantic Web, which has been promoting this idea for almost a decade (and doesn't require centralization). Why?

Here's three anti-SemWeb articles that try to answer your question: http://www.well.com/~doctorow/metacrap.htm http://www.shirky.com/writings/semantic_syllogism.html http://blahsploitation.blogspot.com/2005/09/i-always-figured...

Why is "metacrap" a problem for the semantic web, but not for data-Wikipedia?

The Shirky article is a well-known strawman.

Thanks for the pointer to the last one, I'll read it when I get a chance.

Re: We need a Wikipedia for data

#28
post #14

Two years ago I wrote http://formula1db.com , to teach myself sql. To get the data, I had to screen scrape formula1.com . Once I had the data, learning SQL became a joy. I haven't done much with the data since I built the site, but I am considering open sourcing it. What I don't understand is, why don't sport leagues open source their data. What do they lose? Its a good think that people are so excited about your spo…

Actually, MLB data is fairly close to being open-sourced. Historical data IS open-sourced (though not by MLB itself):

http://baseball1.com/content/view/57/82/ http://retrosheet.org/

Current major and minor league data is available as well, though MLB will crack down on anyone who is trying to make money off of derivative products. Here's where you'll find it, as XML:

http://gdx.mlb.com/components/game/

One can do pretty cool stuff with all of it, and many people have, despite the fact that we can't make money off of it:

http://minorleaguesplits.com/ (my site)

http://baseball.bornbybits.com/2008/pitchers.html (analysis based on detailed pitch speed / break information that MLB started collecting last year.)

Re: We need a Wikipedia for data

#29

Joel had an article on "commoditizing your complements": http://www.joelonsoftware.com/articles/StrategyLetterV.html . Of course we want to commoditize data, to raise the value of hackers (raise the expected return from hacking). On the other side, data companies want to commoditize hackers, to raise the value of data... a process we are sure to resent.

Absolutely, my first thought on reading this was "Why won't these data companies understand they should throw away their business so us coders can make some money?"

Re: We need a Wikipedia for data

#30
post #15

It's so painfully obvious that this is a good idea, and yet people are skeptical of the Semantic Web, which has been promoting this idea for almost a decade (and doesn't require centralization). Why?

Something like this is really what we need, and the thing that would be really revolutionary about the web. Otherwise, even though people talk about the information singularity and such, there isn't a really high useful information signal. Too much useless or mediocre information is worse than useless because it makes everyone more stupid.

But, this is also not only a technology solution. People have to make a choice themselves to filter and promote good information.

Post reply on HN