the fact that you have to encourage these models and tell them that they can solve these problems and warm up on easier problems seems to indicate that there's something to AI pairing above and beyond prompt
One way they improved the hallucination problem was basically training the models to refuse to do or say something if they are not very sure they can do it. As a side effect, they refuse to work on problems they know are extremely hard.