They used to compare to competing models from Anthropic, Google DeepMind, DeepSeek, etc. Seems that now they only compare to their own models. Does this mean that the GPT-series is performing worse than its competitors (given the "code red" at OpenAI)?
OpenAI has never compared their models to models from other labs in their blog post. Open literally any past model launch post to see that.
I see evaluations compared with Claude, Gemini, and Llama there on the GPT 4o post.