Live data from Hacker News

Wikidata: The first new project from Wikimedia Foundation since 2006

meta.wikimedia.org

21–30 of 51 posts

Re: Wikidata: The first new project from Wikimedia Foundation since 2006

#21
post #4

Missing from the FAQ: What's the difference between Freebase and Wikidata?

It looks like the main difference is two-way integration: instead of just scraping data from Wikipedia dumps to produce a structured database (like Freebase and dbpedia do), it's going to store the canonical version of some of the information there, and pull from it to populate the infoboxes. One of the motivations seems to be to keep the data in sync across Wikipedia languages, so an addition or fix propagates to th…

So they are adding an extra layer?...Who said that CS is the science where everything is solved with an extra level of indirection?

Re: Wikidata: The first new project from Wikimedia Foundation since 2006

#22
post #9

tl;dr: spin-off Wikipedia infoboxes into a seperate project with an API, and then use that data to bootstrap an open data project with broader goals. In theory, it's a good idea. It takes an existing useful data source and puts in a form that encourages reuse, and since it solves the bootstrapping problem then it's not obviously doomed to failure like the Semantic Web. I see two potential downsides. My first concern…

Is it even possible to have a database of factual content under CC-BY-SA? This is part of the reason OpenStreetMap is moving to ODbL. Somewhat ironically , since part of the reason is that you can't copyright facts, they didn't just take the existing data under the same theory, but asked everyone to accept the new licence. I wonder what Wikipedia plan to do?

I don't see why you couldn't have a database of facts under CC-BY-SA. You can't copyright individual facts, but you absolutely can copyright a collection of facts as a collection. [1]

I would think the more-pressing problem would be the 'viral' nature of the 'share alike' restriction when it came to API use.

Attribution would also seem to be thorny and difficult to police, but not intractable.

[1] e.g. I can make a phone directory and copyright it. You could take all the data out of my phone directory to make your own directory and that would be fine. But you could not simply make copies of my directory and sell those as your own.

Re: Wikidata: The first new project from Wikimedia Foundation since 2006

#23
post #10

Like others here, it's something I've been thinking about for a number of years. This is an important project, with the potential to eclipse wikipedia, maybe even growing to be the saviour of free software? My reasoning follows. Currently we program computers by giving them a set of instructions on how to achieve a goal. As computers grow more powerful, we will stop giving detailed instructions. Instead, we will writ…

I'm with you to an extent, but why do you think Wikidata in particular will be the missing component and not some other service like Freebase or DBPedia?

Re: Wikidata: The first new project from Wikimedia Foundation since 2006

#24
post #10

Like others here, it's something I've been thinking about for a number of years. This is an important project, with the potential to eclipse wikipedia, maybe even growing to be the saviour of free software? My reasoning follows. Currently we program computers by giving them a set of instructions on how to achieve a goal. As computers grow more powerful, we will stop giving detailed instructions. Instead, we will writ…

"Free software needs Wikidata, to [] avoid being made largely irrelevant by Alpha"

Wolfram Alpha is already completely worthless because it doesn't cite the sources for any of its results. It's basically just a fancy search engine built on top of a garbage dump.

Re: Wikidata: The first new project from Wikimedia Foundation since 2006

#25
post #5

This is actually a startup idea I've had for a while now. It's a great idea in theory, but it's very tricky in practice. Facts have a mysterious way of vanishing if you look closely enough at them, and the raw numbers themselves don't actually tell you anything. The part that's actually interesting is: - The methodology behind the numbers - What we think is most likely the case based on the evidence available - How e…

You couldn't be more right, and I think the key here is: How each fact connects with other facts

If there were no operations, math would just be numbers on their own -- and what fun is that?

The problem is that the relations turn it into the Semantic Web, and after trying and failing to crack that nut for so long, everyone is turned off of it. Which is too bad, because what was failing was the approach. Trying several shipping routes to the New World and failing each time doesn't mean that the New World doesn't exist.

Re: Wikidata: The first new project from Wikimedia Foundation since 2006

#26
post #10

Like others here, it's something I've been thinking about for a number of years. This is an important project, with the potential to eclipse wikipedia, maybe even growing to be the saviour of free software? My reasoning follows. Currently we program computers by giving them a set of instructions on how to achieve a goal. As computers grow more powerful, we will stop giving detailed instructions. Instead, we will writ…

I think one problem is that it's really hard to do structured data in general. Projects that pick a specific domain tend to do it much better, because they have a more tractable problem, can build a community with domain expertise, etc., in ways that Wikipedia will have trouble matching unless they plan to collaborate with those projects and/or pull data from them. For example, I think a structured-data version of Wikipedia artist/album infoboxes is going to have a long way to go to catch up to http://musicbrainz.org/, which has a carefully thought out ontology and years of iteration on that specific problem. Alternatively you can try to do a carefully thought out, consistent schema for all metadata, but the Cyc project shows how hard that is.

I do think that by virtue of breadth Wikipedia's version may become the best data resource in niches that have no specialized structured-data project for them, and it may give other informal-schema, broad-coverage projects like ConceptNet a competitor.

Re: Wikidata: The first new project from Wikimedia Foundation since 2006

#27
post #23
post #10

Like others here, it's something I've been thinking about for a number of years. This is an important project, with the potential to eclipse wikipedia, maybe even growing to be the saviour of free software? My reasoning follows. Currently we program computers by giving them a set of instructions on how to achieve a goal. As computers grow more powerful, we will stop giving detailed instructions. Instead, we will writ…

I'm with you to an extent, but why do you think Wikidata in particular will be the missing component and not some other service like Freebase or DBPedia?

I don't really. Substitute any free body of structured data for Wikidata, or even view them as one body of data, which happens to be spread across multiple servers (and maybe requiring some translation for unification).

Re: Wikidata: The first new project from Wikimedia Foundation since 2006

#28
post #10

Like others here, it's something I've been thinking about for a number of years. This is an important project, with the potential to eclipse wikipedia, maybe even growing to be the saviour of free software? My reasoning follows. Currently we program computers by giving them a set of instructions on how to achieve a goal. As computers grow more powerful, we will stop giving detailed instructions. Instead, we will writ…

"Free software needs Wikidata, to [] avoid being made largely irrelevant by Alpha" Wolfram Alpha is already completely worthless because it doesn't cite the sources for any of its results. It's basically just a fancy search engine built on top of a garbage dump.

http://www.wolframalpha.com/input/?i=how+heavy+is+earth

Click "Source information."

Re: Wikidata: The first new project from Wikimedia Foundation since 2006

#29
post #28

Earlier quoted context omitted.

"Free software needs Wikidata, to [] avoid being made largely irrelevant by Alpha" Wolfram Alpha is already completely worthless because it doesn't cite the sources for any of its results. It's basically just a fancy search engine built on top of a garbage dump.

http://www.wolframalpha.com/input/?i=how+heavy+is+earth Click "Source information."

That isn't a list of references, that's just a list of suggested reading. In fact it's not even guaranteed that the any of the facts on that page come from any of those sources. It's basically just showing a list of books that come up when you Google for the question.

Re: Wikidata: The first new project from Wikimedia Foundation since 2006

#30
post #22

Earlier quoted context omitted.

Is it even possible to have a database of factual content under CC-BY-SA? This is part of the reason OpenStreetMap is moving to ODbL. Somewhat ironically , since part of the reason is that you can't copyright facts, they didn't just take the existing data under the same theory, but asked everyone to accept the new licence. I wonder what Wikipedia plan to do?

I don't see why you couldn't have a database of facts under CC-BY-SA. You can't copyright individual facts, but you absolutely can copyright a collection of facts as a collection. [1] I would think the more-pressing problem would be the 'viral' nature of the 'share alike' restriction when it came to API use. Attribution would also seem to be thorny and difficult to police, but not intractable. [1] e.g. I can make a p…

But being able to legally take all the data out and making your own database (or other thing) with it (which you state is fine) is exactly what makes CC-BY-SA pointless/inapplicable to databases of open data.

See this discussion of why CC-BY-SA is unsuitable for OpenStreetMap (which mentions the case law on phone books you refer to):

http://www.osmfoundation.org/wiki/License/Why_CC_BY-SA_is_Un...

Wikipedia says this on Fiest vs Rural and collections of facts:

"In regard to collections of facts, O'Connor states that copyright can only apply to the creative aspects of collection: the creative choice of what data to include or exclude, the order and style in which the information is presented, etc., but not on the information itself. If Feist were to take the directory and rearrange them it would destroy the copyright owned in the data.

The court ruled that Rural's directory was nothing more than an alphabetic list of all subscribers to its service, which it was required to compile under law, and that no creative expression was involved. The fact that Rural spent considerable time and money collecting the data was irrelevant to copyright law, and Rural's copyright claim was dismissed."

http://en.wikipedia.org/wiki/Feist_v._Rural

Post reply on HN