Earlier quoted context omitted.
I mean 2 years ago they were at about the same place, theres been very little practical gain from gpt4 in my opinion. No matter the model the fundamental failure cases have remained the same.
Reasoning models like o1 or QwQ absolutely destroy 4o in coding, let alone GPT-4 circa 2022.
https://www.youtube.com/live/outcGtbnMuQ?si=oTMA02ns_BJDRS4c...
Advances since then have indeed been remarkable.