Earlier quoted context omitted.
I personally think improvement has been minimal since ChatGPT was released actually.
That's not defensible.
What has improved isn't the models, it's the harnesses.
Give GPT-3.5 a 1M context window and a modern harness, and you won't see any meaningful difference with Opus 5.
It's a bit hard to try with such old models, but for example I use Opus 5 / Fable at work and Sonnet 4.5 at home (because it's free via Amazon Q), and there's absolutely 0 difference in performance. None. Obviously 4.5 is only a year old, not 3, but try with any older model that has a decent context window and you'll get the same results.
In fact I'll go further than this and say that models are currently regressing. Opus 5 is much much worse than Opus 4.6 for example, and it's clear that Anthropic (at least - I don't use OpenAI models much) is just tokenmaxing rather than optimizing for performance.