* This paper called the models directly via the API not the ChatGPT application. This means that changes to the ChatGPT system prompt and other changes to the application aren't a factor here.
* The paper compared two variants each of the chatgpt and GPT4 models. The later variants are obviously different from the earlier variants in some way (likely having been fine-tuned)
* Any given model variant has not changed. You may continue to select the older model variant when using the API if you so wish
Lastly, and this one's my opinion, problems involving arithmetic and mathematics are not a good test of large language models.