Earlier quoted context omitted.
It's certainly RLHFed. All of the logic puzzles I use for evaluation that used to fail months ago now pass no problem and I've even had a hard time modifying them to fail.
This is sort of a bummer because it’s not actually an improvement to the model, but just a patch job to artificially inflate performance. All it does is make true evaluation more difficult. Classic “you get what you measure”.
In that case, just make new problems. If it is being 'patched' to pass specific known problems, then the new ones would fail.
If it is able to answer them, then maybe it is actually analyzing them and working out the solution.
Not sure how you can assume there was no underlying improvement, and these are cases of feeding it the answers.