The paper includes several examples of their modified questions. There has been a substantial jump from o1-preview to o1, so I gave several samples to o1 and o1-pro ( not o1-preview), and current o1s gave the correct answer to those modified problems. SOTA changes fast.
LLM boosters are so tiresome. You hardly did a rigorous evaluation, the set has been public since October and could have easily been added to the training data.
Terence Tao is skilled enough, and he describes O1's math ability is "...roughly on par with a mediocre, but not completely incompetent graduate student" (good discussion at https://news.ycombinator.com/item?id=41540902), and the next iteration O3 just got 25% on his brand new Frontier Math test.
Seeing LLMs as useless is banal, but downplaying their rate of improvement is self-sabotage.