Earlier quoted context omitted.
It's still loss being backproped, but the loss is calculated over a different criteria
Ok that makes a lot of sense. Why do they call it reinforcement learning then? Is it not traditional RE such as Q learning?
The LLM is trained to increase this reward score (or minimize the inverse), which is what makes it RL.