Just as a followup for those interested, after reading the paper and some more critique:
The paper is mostly ok in that it points out model drift and the importance of making sure you use API versions. One good thing is that the full dataset/methodology was published on Github: https://github.com/lchen001/LLMDrift (everyone should do this!) so it was easy to replicate/validate, but there are some issues:
* Simon Boehm stripped the markdown output from the June model output and shows that it actually performs signficantly better than the March update when the Markdown is stripped - 70% correct vs 52% correct. https://twitter.com/Si_Boehm/status/1681801371656536068 - Matei Zaharia (co-author) replies that the point of the paper is that you have to watch out for formatting changes in the LLM, but Matei also announced in the parent tweet "We found big changes including some large decreases in some problem-solving tasks," so which is it? https://twitter.com/matei_zaharia/status/1681467961905926144 - I also think the authors should have been well aware that their prominent Figure 1 would be interpreted as reduced capabilities and that they shouldn't be implying that it is. I'm just going to say it's problematic and leave it at that...
* Narayanan (and colleague Sayash Kappor) published an analysis of the Primality Test https://www.aisnakeoil.com/p/is-gpt-4-getting-worse-over-tim... - basically, the March model didn't do any better than the June model. It's just that one tends to say yes, and one tends to say no, and the Prime factor questions that the authors asked were all "yes." They showed by flipping the questions, suddenly the June update does way better. So one, from a methodology perspective the distribution of yes/no should be equal, but secondly neither model actually has the capability to factor primes, so what's the point of judging if it's correct or not? If the June update for some reason said yes 500 times it would have scored a 100%, but wouldn't mean it actually made a difference on the results. Seems like a pointless test.
* Another thing to note from looking at the code is that they call the API w/ temperature 0.1, not 0.0 - LLM responses are non-deterministic even at 0.0, but I don't know why you'd set it to 0.1 in the first place. The question has been asked: https://github.com/lchen001/LLMDrift/issues/2
I wrote up a more points as well, will just leave a link for those interested: https://fediverse.randomfoo.net/notice/AXsoIewi72IUeXL888