My hunch is that 90-99% of all Jeopardy questions can be answered with information in Wikipedia/Wiktionary, properly understood. So I'd start with Wikipedia: ~30GB uncompressed full article text. Break it into chunks; canonicalize phrasings to be more declarative, and include synonyms/hypernym/hyponym phrasings (via something like WordNet), so that various 'cluesy' ways of saying things still bring up the same candid…
"properly understood" is the whole point. That's the hard problem they are trying to solve.
How to build your own "Watson Jr." in your basement
21–29 of 29 posts
Re: How to build your own "Watson Jr." in your basement
#22Why man? Why? If you want to track clicks on links with javascript, dont trigger the ajax call when I do click in :not(a) ... -_-'
Re: How to build your own "Watson Jr." in your basement
#23The race is on, whoever creates the first reasonably good question-answering machine for demonstration will get their names etched into the sands of time for the next ten thousand years. Get to it!
This industry has the chance to be bigger than Google and Microsoft combined. Every person on the Earth will demand one of these. Those who won't have one will be at a remarkable disadvantage. This is going to turn into a trillion dollar industry.
Re: How to build your own "Watson Jr." in your basement
#24Watson has become an unbelievable marketing tool for IBM.
Has become? From the beginning, Watson's purpose has been to advertise IBM's computers.
Re: How to build your own "Watson Jr." in your basement
#25Earlier quoted context omitted.
Sure, but it does put somewhat of a cap on the amount of reference material you need to import. And a fairly low cap in the tens of GB: Wikipedia/Wiktionary/WordNet/Freebase/J!Archive is probably enough. Beyond that, you want software/heuristics. You might find far more data helpful to initially create that software, but once it's created, the reference material to have at hand can come from a small set of sources.
there is a project on formalization of the human "common sense" with (partly) open source database. take a look: http://en.wikipedia.org/wiki/Cyc
Re: How to build your own "Watson Jr." in your basement
#26My hunch is that 90-99% of all Jeopardy questions can be answered with information in Wikipedia/Wiktionary, properly understood. So I'd start with Wikipedia: ~30GB uncompressed full article text. Break it into chunks; canonicalize phrasings to be more declarative, and include synonyms/hypernym/hyponym phrasings (via something like WordNet), so that various 'cluesy' ways of saying things still bring up the same candid…
"properly understood" is the whole point. That's the hard problem they are trying to solve.
Given that Jeopardy focuses a lot on literature, I'd also throw in gutenburg project books. And probably a newspaper archive going way back, like the NYT.
Re: How to build your own "Watson Jr." in your basement
#27This is one of the great moments in the history of humanity, right up there with the first self-powered flying machine. North Carolina got a licence plate: "FIRST IN FLIGHT". Someone is going to get the credit for open ended question answering machine shortly. Who gets it? The race is on, whoever creates the first reasonably good question-answering machine for demonstration will get their names etched into the sands…
Re: How to build your own "Watson Jr." in your basement
#28Re: How to build your own "Watson Jr." in your basement
#29My hunch is that 90-99% of all Jeopardy questions can be answered with information in Wikipedia/Wiktionary, properly understood. So I'd start with Wikipedia: ~30GB uncompressed full article text. Break it into chunks; canonicalize phrasings to be more declarative, and include synonyms/hypernym/hyponym phrasings (via something like WordNet), so that various 'cluesy' ways of saying things still bring up the same candid…
- The first piece: http://www.slate.com/id/2284678/
- The follow-up, in which I answer reader questions: http://www.slate.com/id/2287705/