OpenAI's models feel 100% nerfed to me at this point. I had it solving incredibly complex problems a few months ago (i.e. write a minimal PDF parser example), but today you will get scolded for asking such a complicated task of it. I think they programmed a classifier layer to detect certain coding tasks and shut it down with canned BS. I like to imagine certain billion/trillion-dollar mega corps had a back-room say…
Do you think a lot of that is scaling pain... like what if they're making cuts to the more expensive reasoning layers to gain more scale. Seems more plausible to me that the teams keeping the lights on have been doing optimization work to save cost and improve speed. The result during those optimizations might not be immediately obvious to the team and then they push deploy and only through anecdotal evidence such as…
In my experience, it's been a mixed bag - had 1 instance recently where it refused to do a bunch of repetitive code, another case where it was willing to tackle a medium complexity problem.