One way they improved the hallucination problem was basically training the models to refuse to do or say something if they are not very sure they can do it. As a side effect, they refuse to work on problems they know are extremely hard.