I really like this line of work and I expect will grow quite substantially over the next few years. Of course, Reinforcement Learning has been around for a long time. Similarly, Q Learning (the core model in this paper) has been around a very long time. What is new is that normally you see these models applied to toy MDP problems with simple dynamics, and linear Q function approximations for fear of non-convergence e…
In fairness, we weren't born with that model; we have to laboriously acquire it over a period of several years. An infant's flailings can look pretty random :-)