I applaud the effort but the problem with RL as a model of learning is in the definition of RL itself. The idea of using "rewards" as a primary learning mechanism and a path to actual cognition is just wrong, full stop. It's a wrong level of abstraction and is too wasteful in energy spent. Looking at it from CogSci perspective it is essentially an offshoot of behaviorism, using a coarse and extremely inefficient mode…
A lot of research in RL is focused on intrinsic motivation and the question of whether we can bootstrap our own 'rewards' from our ability to predict and control the future according to some self-defined goals/hypotheses.