I think a lot of those trick questions outputting stupid stuff can be explained by simple economics. It's just not sustainable for OpenAI to run GPT at the best of its abilities on every request. Their new router is not trying to give you the most accurate answer, but a balance of speed/accuracy/sustainable cost on their side. (kind of) a similar thing happened when 4o came out, they often tinkered with it and the re…
> It's just not sustainable for OpenAI to run GPT at the best of its abilities on every request.
So how do I find out whether the answer to my question was run on the discount hardware, or whether it's actually correct?