Earlier quoted context omitted.
> we don't switch to heavily quantized models That sounded like a press bulletin, so just to let you clarify yourself: Does that mean you may switch to lightly quantized models?
There's almost 0% chance that OpenAI doesn't quantize the model right off the bat. I am willing to bet large amounts of money that OpenAI would never release a model served as fully BF16 in the year of our lord 2026. That would be insane operationally. They're almost certainly doing QAT to FP4 for FFN, and a similar or slightly larger quant for attention tensors.
Arena AI Model ELO History
11–20 of 63 posts
Re: Arena AI Model ELO History
#12Re: Arena AI Model ELO History
#13I hope to see the other labs can bring back competition soon!
Re: Arena AI Model ELO History
#14Re: Arena AI Model ELO History
#15FYI, Elo isn't an acronym - it's a person's name. No need to capitalize it as ELO.
Re: Arena AI Model ELO History
#16The interesting thing I find is how Anthropic has been more consistently improving over time in the last few years, that allows it to catchup and surpass OpenAI and Google. The latter two have pretty much plateau over the last year or so. GPT 5.5 is somehow not moving the needle at all. I hope to see the other labs can bring back competition soon!
Re: Arena AI Model ELO History
#17Re: Arena AI Model ELO History
#18Honestly, in my opinion, GPT-5.5 Codex doesn't just crush Claude Code 4.7 opus —it's writing code at a level so advanced that I sometimes struggle to even fully comprehend it. Even when navigating fairly massive codebases spanning four different languages and regions (US, China, Korea, and Japan), Codex's performance is simply overwhelming.
How would we even go about properly measuring and benchmarking the Elo for autonomous agents like this?
Re: Arena AI Model ELO History
#19The interesting thing I find is how Anthropic has been more consistently improving over time in the last few years, that allows it to catchup and surpass OpenAI and Google. The latter two have pretty much plateau over the last year or so. GPT 5.5 is somehow not moving the needle at all. I hope to see the other labs can bring back competition soon!
Gpt 5.5 is quite a big leap, it's a lot better than opus 4.7 for agentic coding
Re: Arena AI Model ELO History
#20This is great, but personally, I really wish we had an Elo leaderboard specifically for the quality of coding agents. Honestly, in my opinion, GPT-5.5 Codex doesn't just crush Claude Code 4.7 opus —it's writing code at a level so advanced that I sometimes struggle to even fully comprehend it. Even when navigating fairly massive codebases spanning four different languages and regions (US, China, Korea, and Japan), Cod…