One thing I wonder about hallucinations, is that it seems on the surface that it is an easy problem for RLVR to target. Since you're already generating enormous amounts of reasoning traces which are verified by correct answers, just have "don't know" as an option as a valid answer, and on problems where none of the thousands of reasoning traces led to a correct answer, just promote the traces that led to the "don't k…
But if an LLM says "I don't know" should you pay for the tokens?
We can rank them based on how much they know and people will gravitate towards those that do know more.
It's a market after all.