Earlier quoted context omitted.
To me, this is a tell of human-involvement in the model solution. There is no reason why machines would do badly on exactly the problem which humans do badly as well - without humans prodding the machine towards a solution. Also, there is no reason why machines could not produce a partial or wrong answer to problem 6 which seems like survivor bias to me. ie, that only correct solutions were cherrypicked.
Maybe it’s a hint that our current training techniques can create models comparable to the best humans in a given subject, but that’s the limit.
Noam Brown: 'This result is brand new, using recently developed techniques. It was a surprise even to many researchers at OpenAI.'
So your thesis is that these new techniques - which just produced unexpected breakthroughs - represent some kind of ceiling? That's an impressive level of confidence about the limits of methods we apparently just invented in a field which seems to, if anything, be accelerating.