Earlier quoted context omitted.
It hasn't read that riddle because it is a modified version. The model would in fact solve this trivially if it _didn't_ see the original in its training. That's the entire trick.
Sure but the parent was praising the model for recognizing that it was a riddle in the first place: > Whereas o1, at the very outset smelled out that it is a riddle That doesn't seem very impressive since it's (an adaptation of) a famous riddle The fact that it also gets it wrong after reasoning about it for a long time doesn't make it better of course
If you are tricked about the nature of the problem at the outset, then all reasoning does is drive you further in the wrong direction, making you solve the wrong problem.