In these experiments the neurons don't know they are firing a gun and chasing demons. The task is to produce outputs (movement) to centre an input and then produce an output (shoot). Then the cycle repeats. Typically the input it receives is increasingly regular as it does better. And irregular to mark a failure which the neurons learn to avoid. It has no idea that the data can also be rendered as a Doom game.
Does that mean human brains have neuron "training" feedback?
Most fundamentally come down to the neurones connections and behaviour can be modelled as minimizing prediction error in their stimulus.
Dopamine encoding an outcome compared to a baseline expectation is very close to tempeoal difference learning.
This is basically what slot machines exploit with their variable reward schemes.