More interesting info on Maluuba's 2 actual, recently-released datasets: http://datasets.maluuba.com/ Their "News QA dataset" contains 120k Q&As collected from CNN articles: Documents are CNN news articles. Questions are written by human users in natural language. Answers may be multiword passages of the source text. Questions may be unanswerable. NewsQA is collected using a 3-stage, siloed process. Questioners see o…
"Frame tracking" and adaptive slot filling without pre-canning is actually a very active area of research - the goal is to provide not just an Eliza style infinite conversational ability, but to be able to reason about and get the user to a specific outcome in a real world use case by prompting them for information in a semi supervised manner.
Being able to do it via a series of differentiable (calculus wise) functions is a fundamental improvement in achieving convergence between statistical systems and logical reasoning (which harks back to "classical" AI problems such as planning). Microsoft has some very interesting research here and papers like these [1][2]. Maluuba has some good stuff here as well. [3][4]
I doubt the company was acquired for their datasets. They've been a well respected AI company with some top notch researchers, and I think their area of research gels well with MSR's own NLP research.
[1] https://arxiv.org/abs/1606.01269
[2] https://www.microsoft.com/en-us/research/publication/unsuper...