Just waiting on what Terrance Tao thinks of the new o3 intern
"Assessment of FrontierMath difficulty
All four mathematicians characterized the research problems in the FrontierMath benchmark as exceptionally challenging, noting that the most difficult questions require deep domain expertise and significant time investment. For example, referring to a selection of several questions from the dataset, Tao remarked, "These are extremely challenging. I think that in the near term basically the only way to solve them, short of having a real domain expert in the area, is by a combination of a semi-expert like a graduate student in a related field, maybe paired with some combination of a modern AI and lots of other algebra packages...”
However, some mathematicians pointed out that the numerical format of the questions feels somewhat contrived. Borcherds, in particular, mentioned that the benchmark problems “aren’t quite the same as coming up with original proofs.”
It sounds like it will be able to crack some hard math problems, but not actually do mathematics. Which makes sense to me, given how these oN models are being trained. Synthetic data is bound to be in some (more or less obvious) way contrieved, since it's not like we can just generate tons of new/natural/original mathematical results to train the model on.