The interesting thing I find is how Anthropic has been more consistently improving over time in the last few years, that allows it to catchup and surpass OpenAI and Google. The latter two have pretty much plateau over the last year or so. GPT 5.5 is somehow not moving the needle at all. I hope to see the other labs can bring back competition soon!
Gpt 5.5 is quite a big leap, it's a lot better than opus 4.7 for agentic coding
Arena AI Model ELO History
31–40 of 63 posts
Re: Arena AI Model ELO History
#32For what it's worth, I work at OpenAI and I can guarantee you that we don't switch to heavily quantized models or otherwise nerf them when we're under high load. It's true that the product experience can change over time - we're frequently tweaking ChatGPT & Codex with the intention of making them better - but we don't pull any nefarious time-of-day shenanigans or similar. You should get what you pay for.
Re: Arena AI Model ELO History
#33Re: Arena AI Model ELO History
#34Re: Arena AI Model ELO History
#35> the slow performance decays the decays are just more capable other models entering the population, making all prior models lose more frequently
No, that is not how ELO scores work.
Re: Arena AI Model ELO History
#36Re: Arena AI Model ELO History
#37Re: Arena AI Model ELO History
#38> the slow performance decays the decays are just more capable other models entering the population, making all prior models lose more frequently
No, that is not how ELO scores work.
Re: Arena AI Model ELO History
#39FYI, Elo isn't an acronym - it's a person's name. No need to capitalize it as ELO.
Re: Arena AI Model ELO History
#40Seems like Chinese labs are the only ones that are trustworthy (at least when it gets to this specific issue). This feels so ironic haha
Novita's has occassional problem counting white space. DeepSeek hosted does not.
No idea why.