So there are plenty of reasons to be skeptical of OpenAI, and anyone looking at my comment history will see that I'm not shy about calling them out on real things (the F-you name, the crazy "is this even legal?" structure, the continued insistence this isn't for the personal benefit of the people involved, the lobbying, the arms-length affiliation with the dystopian biometrics thing, the list goes on).
But this and a few other recent headlines around the ChatGPT/Instruct frontend to the 3.5-turbo and 4-series models having stability issues isn't a reasonable thing to get on their case about: what company operating at that kind of scale doesn't have stability/quality issues? If anything it's really friggin impressive how little operational bleed they've had on what is arguably the toughest ops problem a major Internet property has ever faced at launch. How the hell do you even smoke test something like that?
Now to the extent that they say or imply that we're not all getting A/B-tested at a minimum, and more likely explore-exploit bandited against inference costs, that's almost certainly horseshit: they clearly started selling the 4-series at least and probably all of them at a loss on the inference, and they've clearly been playing with ways to control costs (Rich Uncle Nadella's patience with loss-leaders is most likely finite, being as he's not running a charity and all).
But as someone who finds a lot of this loathsome on the social and business side, I'm not going to knock an engineering and ops group for tripping up here and there on a problem that hard. That's not fair.