I feel very comfortable saying, as a mathematician, that the ability to solve grade school maths problems would not be at all a predictor of ability to solve real mathematical problems at a research level. The reason LLMs fail at solving mathematical problems is because: 1) they are terrible at arithmetic, 2) they are terrible at algebra, but most importantly, 3) they are terrible at complex reasoning (more specifica…
2. Often, research is made on toy models and if this would be such a model, acing grade school problems (as per the article) would be quite impressive to say the least as this ability simply isn't emergent early in current LLM's.
What I think might have happened here is a step forward in AI capacity that has surprised researchers not because it is able to do things it couldn't at all do before, but how _early_ it is able to do so.