Like others have said, this seems to be a generalization of the problem of reinforcement learning. For a good introduction to the subject, check out Reinforcement Learning by Sutton & Barto [1]. After reading the first few chapters, you'll be able to understand most of that equation. [1] http://webdocs.cs.ualberta.ca/~sutton/book/the-book.html
You just use the rewards to optimize the function that tells you what's the predicted reward at any stage for a given action. And then take those best acions