All of the tests are one-shot questions and answers. Where I have found GPT-4 to be degrading significantly is with sustained discussion about technical topics. It starts forgetting important parts of the discussion almost straight away, long before the size of the context window becomes a factor. This wasn’t the case when it was new.
So, I asked a follow up in the next prompt and in the next response it was wildly off-base and it's response made no sense and it had hallucinated all these functions into my code that had no business there. 2 prompts from me, 1 response from gpt and then it's second response it is completely lost.