It has become common knowledge that GPT4 (and also 3.5) have problems with deterministic outputs (even at T=0). So what we're seeing here is just the effect of random sampling, not any actual change to the model itself. If you scroll down, you'll see other close attempts by the exact same model that could already be counted as a win depending on who you ask. Edit: This comment section is a super fascinating case stud…
1. A change in the current model number (eg. gpt-3.5-turbo-0613)
2. On ChatGPT UI, the date at the bottom (eg. August 2023)
So it isn’t correct to say “it is incredibly obvious that nothing has happened”. Not that obvious to me.
A bit like how you can never tell for sure if Coca Cola has tweaked their formula, or McDonalds has changed the recipe for its signature sauce. Only in this case, the model number going up or the date becoming more recent leads credence to something having changed.