Live data from Hacker News

We need a Wikipedia for data

bret.appspot.com

11–20 of 56 posts

Re: We need a Wikipedia for data

#11
"No one really wants factual data accuracy and completeness to be their competitive advantage"

Some people actually do build their business on this. It's not the most sustainable business model, but in today's world, good data can be - and is - a competitive advantage. In an ideal world, should it be that way? Maybe not. I think that's what this post is getting at.

Re: We need a Wikipedia for data

#13
post #6

It's like freebase all right, but the article has something right about adoption. He points out some big company would have to donate a great starting dataset to drive adoption. I think this is one problem with Freebase. Another problem I see is the structuring of the data, it is a hurdle to sharing. Finally the last problem I see is Data is currently very much seen as a competitive advantage. When Google introduce f…

It's like freebase all right, but the article has something right about adoption. He points out some big company would have to donate a great starting dataset to drive adoption.

http://news.ycombinator.com/item?id=157966

Re: We need a Wikipedia for data

#14
Two years ago I wrote http://formula1db.com , to teach myself sql. To get the data, I had to screen scrape formula1.com . Once I had the data, learning SQL became a joy. I haven't done much with the data since I built the site, but I am considering open sourcing it.

What I don't understand is, why don't sport leagues open source their data. What do they lose? Its a good think that people are so excited about your sport, that they build custom apps based on it. Sadly sport leagues don't seem to get it, I remember the MLB cracking down on a fan generated datatbase of baseball statistics a while ago.

Re: We need a Wikipedia for data

#15
It's so painfully obvious that this is a good idea, and yet people are skeptical of the Semantic Web, which has been promoting this idea for almost a decade (and doesn't require centralization). Why?

Re: We need a Wikipedia for data

#16
post #10

Bret is spot on. Open data would unlock a vast amount of wealth. Other suggestions: A collection of various kinds of texts, translated into 20-30 different languages. 100 million words (per language) would be fine. 20 minutes of text, spoken in thousands of different voices/accents.

An open translation dictionary is a fantastic idea. You would have to be careful to clarify the context. A word in one language can often translate to several words in another language, for example. But I think it's do-able.

And I think you meant 100 thousand words per language. :-)

Re: We need a Wikipedia for data

#18
I've thought about this before... Wikipedia covers unstructured data, but we need something to be its analog for structured data. Others have pointed out freebase (which kicks ass IMO), but there's also swivel (http://www.swivel.com/). Swivel's correctness model, however, doesn't seem to be the same open idea that freebase is. It seems more to be focused on data "authorities" and "official" data providers as opposed to just an "accuracy of the masses" system that the more open sites rely on.

Re: We need a Wikipedia for data

#19
post #14

Two years ago I wrote http://formula1db.com , to teach myself sql. To get the data, I had to screen scrape formula1.com . Once I had the data, learning SQL became a joy. I haven't done much with the data since I built the site, but I am considering open sourcing it. What I don't understand is, why don't sport leagues open source their data. What do they lose? Its a good think that people are so excited about your spo…

They think that they can't sell the data if they open source it. They probably also have control issues.

They think that a popular application using "their data" is necessarily lucrative and they think that they should get a huge hunk of that money.

Post reply on HN