These are high school level only in the sense of assumed background knowledge, they are extremely difficult. Professional mathematicians would not get this level of performance, unless they have a background in IMO themselves. This doesn’t mean that the model is better than them in math, just that mathematicians specialize in extending the frontier of math. The answers are not in the training data. This is not a mode…
E.g here: https://pbs.twimg.com/media/GwLtrPeWIAUMDYI.png?name=orig
Frankly it looks to me like it's using an AlphaProof style system, going between natural language and Lean/etc. Of course OpenAI will not tell us any of this.