Live data from Hacker News

Arena AI Model ELO History

mayerwin.github.io

31–40 of 63 posts

Re: Arena AI Model ELO History

#31
post #16
post #13

The interesting thing I find is how Anthropic has been more consistently improving over time in the last few years, that allows it to catchup and surpass OpenAI and Google. The latter two have pretty much plateau over the last year or so. GPT 5.5 is somehow not moving the needle at all. I hope to see the other labs can bring back competition soon!

Gpt 5.5 is quite a big leap, it's a lot better than opus 4.7 for agentic coding

Better in what ways? I'm just curious about your experience.

Re: Arena AI Model ELO History

#32

For what it's worth, I work at OpenAI and I can guarantee you that we don't switch to heavily quantized models or otherwise nerf them when we're under high load. It's true that the product experience can change over time - we're frequently tweaking ChatGPT & Codex with the intention of making them better - but we don't pull any nefarious time-of-day shenanigans or similar. You should get what you pay for.

its very interesting to see that this only happens to American companies. What gives?

Re: Arena AI Model ELO History

#35
post #29
post #2

> the slow performance decays the decays are just more capable other models entering the population, making all prior models lose more frequently

No, that is not how ELO scores work.

It depends what you use as an anchor. If the anchor is a fixed model, you’re right. If the anchor is updated to a better model over time, then the elo of historical models degrades, right?

Re: Arena AI Model ELO History

#38
post #29
post #2

> the slow performance decays the decays are just more capable other models entering the population, making all prior models lose more frequently

No, that is not how ELO scores work.

As far as I understand, this is exactly how ELO scores work. If a more capable show up and starts beating all the other models, it literally takes ELO points from everyone else.

https://en.wikipedia.org/wiki/Elo_rating_system

Re: Arena AI Model ELO History

#40
post #30

Seems like Chinese labs are the only ones that are trustworthy (at least when it gets to this specific issue). This feels so ironic haha

I am using novita-hosted DeepSeek V4 (Flash) for work and DeepSeek API for personal projects.

Novita's has occassional problem counting white space. DeepSeek hosted does not.

No idea why.

Post reply on HN