Live data from Hacker News

Wikipedia is now drawing facts from the Wikidata repository

gigaom.com

51–60 of 82 posts

Re: Wikipedia is now drawing facts from the Wikidata repository

#51

I'm surprised no one (AFAICT) has attempted a family tree of all humans. There is an obvious demand for this information because many commercial services exist. Users are paying to upload their personal genealogical data to proprietary for-profit silos. Yet this data would be much more productive in an open system with user data from all services. You could seed the database with famous people's family trees from Wik…

The genealogists are always happy to share their data. It's part of the culture.

There's no open source genealogy programmers because the young nerds that do open source dont care about genealogy.

Re: Wikipedia is now drawing facts from the Wikidata repository

#52
post #17

Earlier quoted context omitted.

This is a good example for Paul Graham's "middlebrow dismissal" data set. Generic negative comment about the subject of the article that could be copy-pasted into any article matching a "Wikipedia" string search by a bot.

I see the brow a great deal lower than you see it. Middlebrow is unduly dignified.

Your comment is itself a negative, middlebrow dismissal. The OPs negativity is middlebrow precisely because it's accurate, but it's beating a dead horse. This means it gets up voted repeatedly, whereas a genuinely low brow comment like "FUK U GAY WIKIPEDIA" would be quickly down voted.

Re: Wikipedia is now drawing facts from the Wikidata repository

#53
post #38
post #25

Earlier quoted context omitted.

The presence of an article on obsolete BBS software (for instance), somewhere in the wikipedia encyclopedia, does not intersect with anyones use-case that isn't already searching for it. Wikipedia isn't one big giant article, where the existence of marginally noteworthy elements distracts you or diverts your attention. The articles that get deleted by the little notability hitlers are contributions with no cost incur…

> The articles that get deleted by the little notability hitlers are contributions with no cost incurred by anyone, save the trivial storage space. Godwin, and wrong on the merits, as well. The stuff that gets deleted is the stuff that can't be verified, which means there's no way to fact-check it. It isn't about storage space or clutter: It's about not having stuff in there that can't be verified.

It's inaccurate to say that wikipedia is fact checked or verified. The existence of a citation does not imply the existence of a fact check or a verification, or verifiability. Even when citations are high quality, the info they're supporting is still usually unverifiable due to wikipedia's disconnection from the community of experts.

Wikipedia is nothing more than the biggest plagiarist / content farm on the Internet. It isn't scrutinized because it has been grandfathered in.

Wikipedia is the ebaumsworld of information. Completely unreliable. Steals credit, traffic, royalties from the content creators. Policies focused on self preservation rather than serving a public good or respecting creators.

Re: Wikipedia is now drawing facts from the Wikidata repository

#54
post #23

Earlier quoted context omitted.

Err on which mission exactly? http://en.wikipedia.org/wiki/List_of_Pok%C3%A9mon

To map all human knowledge.

Isn't non-notable knowledge still knowledge? They are only interested in a small slice of human knowledge by their own sated goals.

Re: Wikipedia is now drawing facts from the Wikidata repository

#55
post #10
post #6

It's too bad that all of these grand visions and truly positive developments for humanity are locked inside of an exclusionary, elitist organization that is governed by deletionism and will take any contribution you might possibly make and CRUSH IT LIKE A BUG.

That's a bit harsh. It's true the admins take notability pretty seriously, but to be fair, so does Encyclopaedia Britannica. Wikipedia folks freely admit that they're elitist in that regard, but I don't think that's necessarily a bad thing. They set it to that level for the sake of maintaining quality and relevancy (which is a measure of notability) to an acceptable degree. The foundation's ultimate goal is to make a…

Britannica actually accepts user generated content. They are superior to wikipedia, however, in that they have credible experts vet the content.

The complaint about wikipedia is that it fails to live up to its ideals. If you had ever been a wikipedia editor you wouldn't talk the way you do. Even as a domain expert it's intolerable to deal with the wikipedia amateur hour culture. It's far more rewarding to submit your articles to Encyclopedia Britannica.

Re: Wikipedia is now drawing facts from the Wikidata repository

#56
post #4

OT: I was using Wikipedia the other day and it occurred to me how primitive it is to have all the inner links to other Wikipedia articles defined manually, surely these should have been automated by now (i.e., marking a word or two would link you to the relevant article).

The disambiguation could be challenging. (Does "The Sun" link to our closest star or the tabloid in the UK.) It would be a fun project to try and determine the correct link based on the context.

This reminds me of an article I read about IBM's Watson. As a demo in front of a lot of people, a researcher was going through a stack of journals and feeding in data about anthrax. Most of the data was about animals, but the researcher was asking Watson to extrapolate possible effects on people. Watson responded "I assume by people you mean humans, and not People magazine."

Edit: found the link (PDF) http://www.cs.mtu.edu/~nilufer/classes/cs5811/2003-fall/hilt... Here's the actual quote:

  “Do you mean Anthrax (the heavy-metal band),
  anthrax (the bacterium) or anthrax (the disease)?”

  “The bacterium,” was the typed answer, followed by 
  the instruction, “Comment on its toxicity to people.”

  “I assume you mean people (homo sapiens),” the
  system responded, reasoning, as it informed its
  programmer, that asking about People magazine
  “would not make sense.”

Re: Wikipedia is now drawing facts from the Wikidata repository

#57

I'm surprised no one (AFAICT) has attempted a family tree of all humans. There is an obvious demand for this information because many commercial services exist. Users are paying to upload their personal genealogical data to proprietary for-profit silos. Yet this data would be much more productive in an open system with user data from all services. You could seed the database with famous people's family trees from Wik…

The genealogists are always happy to share their data. It's part of the culture. There's no open source genealogy programmers because the young nerds that do open source dont care about genealogy.

I care about genealogy! Of course, it's debatable if I'm considered "young" since I'm 30. I think maybe there would be privacy issues and personality rights to photographs that would get in the way of such a big family tree.

Not to mention a security issue as well since most financial institutions ask for your mother's maiden name as a security question.

Re: Wikipedia is now drawing facts from the Wikidata repository

#58
post #46

Earlier quoted context omitted.

> I'm only talking about auto-generated links that can't be clearly disambiguated by the system. But they could be disambiguated by humans, which is my point. Humans understand context.

Sure, and when humans create links, they should continue to create them just like they do not. I'm picturing an "auto linkifier" that creates links that no human has gotten around to creating yet. Whether or not something like that would be a net win for Wikipedia is up for debate I guess. That said, I think they already do have a bot that can do at least a limited amount of auto-linkification, but I can't swear to i…

I guess an added bonus of what you're suggesting is that the correct link could be crowdsourced; if the system kept track of which of the options users clicked on, it could figure out pretty quickly which one is correct.

Re: Wikipedia is now drawing facts from the Wikidata repository

#59
Well, this discussion is quickly moving toward the shortcomings of Wikipedia.

There's plenty of opinion of the deletion-happy (if that's the operative term) policy of the admins, especially as it pertains to notability, and I think this is a common complaint : http://www.highprogrammer.com/alan/rants/wikipedia-delete.ht...

I do something that may seem like trying to fix a leaky dam with chewing gum, but every time I see an article nominated for deletion, I copy the source to a private wiki I'm running. Sometimes, the deletion nomination goes away, other times it does get deleted, but at least then, I have a copy.

But we also have to keep in mind that there are plenty of other resources on the web specifically aimed at niche resources. E.G. Wikia. I've lost count of how many comic book related articles were deleted only to show up on the DC or Marvel wikis. ( http://dc.wikia.com/ http://marvel.wikia.com/ )

Likewise, it's not unreasonable that a lot of content gets missed by the editors, since it's a big place and most editors have one or two areas that they focus on. As the editing guidelines state, when in doubt engage in dispute resolution, not edit wars. If you can make a good case for why an article should be there in the first place, be persuasive in the talk pages.

So a few pointers for people getting angry at Wikipedia:

First, ask whether the article is a good fit for the wiki. Can it go on a blog or a niche wiki (like a dedicated Wikia) instead? The Pokémon example is a bit extreme, but I think many of those pages may get deleted or merged. There's always the Pokémon Wikia : http://pokemon.wikia.com

Second, notability is a very tricky thing. Reputable sources may be even trickier. Rather than debating notability, focus on reputable sources (since that's the biggest hiccup for references). If you can link NY Times articles or BBC or some other news source, rather than just community sites, other blogs (depending on popularity) etc... you'll have a better chance of getting the article/section through and staying there.

Third, well written articles have a better chance of surviving than those that give off the 2-3 paragraph stub vibe. The more reputable citations and best content structure you can give, the better the chances an article survives (this may partly explain the Pokémon pages too). If you have trivia, try to merge it into the content body more, rather than list at the end; that feels tacked on and superfluous, if not directly supporting the main content.

Fourth, try to be a bit more empathetic to the goals of the Wiki while being objective to the subject when asserting your views (especially on controversial articles). How you word things is a big hint as to whether that content will remain or get scrubbed the next day/hour/minute.

Now, I hope can go back to discussing Wikidata and how Wikipedia and everyone else will benefit.

Post reply on HN