Live data from Hacker News

We need a Wikipedia for data

bret.appspot.com

31–40 of 56 posts

Re: We need a Wikipedia for data

#31
post #10

Bret is spot on. Open data would unlock a vast amount of wealth. Other suggestions: A collection of various kinds of texts, translated into 20-30 different languages. 100 million words (per language) would be fine. 20 minutes of text, spoken in thousands of different voices/accents.

An open translation dictionary is a fantastic idea. You would have to be careful to clarify the context. A word in one language can often translate to several words in another language, for example. But I think it's do-able. And I think you meant 100 thousand words per language. :-)

I don't mean a dictionary (also a good idea), I meant texts: articles, novels, blog posts, transcripts of conversations, etc.

As for dictionaries, there's wiktionary, but it's broken because it's based on words, not meanings, so you'd need 30*29 translations for each word.

Mmm, maybe I should do it...

Re: We need a Wikipedia for data

#33
post #24

Earlier quoted context omitted.

Here's three anti-SemWeb articles that try to answer your question: http://www.well.com/~doctorow/metacrap.htm http://www.shirky.com/writings/semantic_syllogism.html http://blahsploitation.blogspot.com/2005/09/i-always-figured...

Why is "metacrap" a problem for the semantic web, but not for data-Wikipedia? The Shirky article is a well-known strawman. Thanks for the pointer to the last one, I'll read it when I get a chance.

I really shouldn't be doing this...

> Why is "metacrap" a problem for the semantic web, but not for data-Wikipedia?

Because Wikipedia is centralized, and the SemWeb isn't.

> The Shirky article is a well-known strawman.

DH3. Contradiction

Re: We need a Wikipedia for data

#34
post #24

Earlier quoted context omitted.

Why is "metacrap" a problem for the semantic web, but not for data-Wikipedia? The Shirky article is a well-known strawman. Thanks for the pointer to the last one, I'll read it when I get a chance.

I really shouldn't be doing this... > Why is "metacrap" a problem for the semantic web, but not for data-Wikipedia? Because Wikipedia is centralized, and the SemWeb isn't. > The Shirky article is a well-known strawman. DH3. Contradiction

I don't want to turn this into a huge debate either, but those articles (and uncritical readings of them) have set the web back years.

> Because Wikipedia is centralized, and the SemWeb isn't.

If data-Wikipedia and a television station are both publishing data about when your favorite show is on that station, who are you more likely to believe?

Obviously you need to be careful about where your data comes from, but a single centralized source is not necessarily more trustworthy than many carefully selected sources.

Blind crawling isn't (and will probably not be) the norm for data collection on the semantic web.

> DH3. Contradiction

Heh, got me there.

Shirky's thesis is based on the idea that making inferences from data is the ultimate purpose of the semantic web.

But linked, machine-readable data--that is, the semantic web--is useful even if inferencing is useless. I don't think this is a claim that needs evidence, it should be fairly obvious.

Shirky's article's portrayal of the semantic web has little to do with the real thing. Here's a much broader debunking of it: http://www.poorbuthappy.com/ease/semantic/

Re: We need a Wikipedia for data

#36
post #5

Isn't Metaweb's Freebase ( http://www.freebase.com/ ) doing this? Check it out if you have not seen this before, these guys are Rockstars!!

Exactly my thought. Freebase is awesome, I just wish they didn't use JSON.

I like them more because they use JSON.

Re: We need a Wikipedia for data

#37
There's largely no such thing as "closed source" data. Many of the restrictions people claim on publicly distributed data are bogus: you cannot claim copyright on a comprehensive collection of facts. http://www.iusmentis.com/databases/us/ http://blog.infochimps.org/2008/04/02/good-neighbors-and-ope... I don't think baseball is cracking down on people making money on this, unless they infringed their (quite reasonable) hot news claims to the real-time data.

Baseball is the leading example of why giving away most of your data is the best use of it. The sport of baseball -- the way it's played on the field, the way players are scouted and trained, and the way it's enjoyed as a fan (Moneyball? Fantasy Sports?) -- have been revolutionized by amateurs making use of free open data.

If you give out the great bulk of your data, people will be enhancing it with metadata, building tools on top of it, and most importantly connecting it to the rest of humanity's knowledge store and mining it for connections you'd have never conceived. Giving out "up to last month" or "daily intervals" will grow sharply the market for "real time" or "second-by-second". Baseball's mission statement concerns bats, bases, butts and seats -- not visualizing correlations among heterogeneous data stores. By releasing their data for free they let the smartest people in the world have the opportunity to perform that second task for free.

We're about to enter the age of ubiquitous information. Drawing these data stores into open formats, making them discoverable, and interconnecting them across knowledge domains presents explosive opportunities. But who will own this data and what access will they allow? If you want to help ensure that the answer is 'everyone' and 'all of it', come join the http://infochimps.org project, a free open community effort to build an Allmanac of everything.

Re: We need a Wikipedia for data

#38
My degree thesis was on this subject, except this...

It stores the data as English text, parseable into Prolog predicates (ie, key/value data chunks), and you have a special editor which shows you the predicates generated by the English you just wrote.

Does this sound useful to anyone? I'd love to pick up on it.

The parser for English I wrote is in Lisp, and works very well (it's recursive, etc).

And of course, you don't have to make a new database, you just do the usual Wiki editing and just make sure the text you enter is parseable by the latest parser.

Re: We need a Wikipedia for data

#39
post #13
post #6

It's like freebase all right, but the article has something right about adoption. He points out some big company would have to donate a great starting dataset to drive adoption. I think this is one problem with Freebase. Another problem I see is the structuring of the data, it is a hurdle to sharing. Finally the last problem I see is Data is currently very much seen as a competitive advantage. When Google introduce f…

It's like freebase all right, but the article has something right about adoption. He points out some big company would have to donate a great starting dataset to drive adoption. http://news.ycombinator.com/item?id=157966

It's a start for sure and it's great that freebase does this. But what I refer to is more like Google gives all their business location listings to Freebase. I don't think this will happen, but maybe something of less value while being still significant.

Re: We need a Wikipedia for data

#40
Finding high quality data is very expensive, and very well-guarded with license restrictions once you do cough up for it..

Seems like factual data will inevitably gravitate toward free, though. The sooner the better.

Post reply on HN