Checking the arithmetic in every paper published seems like an good use case for LLMs. Has someone built a better version than uploading a PDF to ChatGPT and asking it to check the arithmetic?
LLM's are why we're in this mess, they can't do math or count r's
ChatGPT 5.2 has recently been churning through unsolved Erdös problems.
I think right now one is partially validated by a pro and the other one I know of is "ai-solved" but not verified. As in: we're the ones who can't quite keep up.
https://arxiv.org/abs/2601.07421
And the only reason they can't count Rs is that we don't show them Rs due to a performance optimization.