>No evidence of GPT-4 getting worse over time
I read the the article you linked to, and either you're putting words in their mouth and/or making claims that go way beyond the conclusions presented in the article.
The authors choose to distinguish between "capability degradation" and "drastic behavior change." Now, almost no one I've witnessed complaining about past and present results of ChatGPT care one bit about making such a distinction. They often simply complain that their previously good result producing prompts are now spewing garbage.
The authors explicitly admit that "the kind of fine tuning that LLMs regularly undergo can have unintended effects, including drastic behavior changes on some tasks." i.e. Fine tuning may in fact be causing the worse results.
Furthermore, the authors merely make an additional point that such a drastic behavior change shouldn't be assumed to be a "capability degradation." Meaning that, theoretically, if you toy around with your past prompt enough you should be able to eventually find an alternative version of the original prompt that outputs the past result you were looking for. Again, none of the complainers that I've come across really give a shit about this theoretical conclusion. They only care about past prompts now giving different results. In the article, the authors freely admit that this "drastic behavior change" could be objectively occurring due to fine-tuning.
To further prevent readers from making sweeping generalizations from their comments, the author's explicitly state that OpenAI makes it pretty damn difficult to do reproducible research on the LLMs in question, so you should not come to any hasty conclusions regardless ("As we have written before, this underscores how hard it is to do reproducible research that uses these APIs, or to build reliable products on top of them.")
Well, "building reliable products/processes on top of ChatGPT" is all the complainers care about anyway, so we're back to square one. Far from disproving the claim that ChatGPT is becoming more unreliable and producing worse results for the complainers use cases, the authors are in fact saying that their grievances may be justified for the only practical criteria the naysayers actually give a shit about.