ChatGPT 4's math abilities are really hit or miss. I asked it to act as a "Socratic tutor" for the de-biasing terms in exponential moving averages. And it walked me through the problem, and it actually pointed out errors in my derivation (even though I was using a half-baked mix of Unicode and LaTeX). That was honestly pretty impressive.
On the other hand, GPT 3.5 is painfully bad at basic arithmetic. Apparently it can add long numbers, but you actually need to explain the addition algorithm and give it some examples. Then it can walk through the process step by step, just like an elementary school student.
In a lot of cases, I suspect ChatGPT 4 is pulling from a tiny corner of the training distribution. It has seen most of the web, and it doesn't need to see things often to talk about them. I've seen it successfully summarize the contents of obscure personal blogs, or of a very minor news story that was current for a week in 2009.
So it wouldn't surprise me if GPT 4 performs well on tasks that are mentioned on a half dozen obscure web pages, or something like that.