Live data from Hacker News

Arena AI Model ELO History

mayerwin.github.io

11–20 of 63 posts

Re: Arena AI Model ELO History

#11
post #9
post #7

Earlier quoted context omitted.

> we don't switch to heavily quantized models That sounded like a press bulletin, so just to let you clarify yourself: Does that mean you may switch to lightly quantized models?

There's almost 0% chance that OpenAI doesn't quantize the model right off the bat. I am willing to bet large amounts of money that OpenAI would never release a model served as fully BF16 in the year of our lord 2026. That would be insane operationally. They're almost certainly doing QAT to FP4 for FFN, and a similar or slightly larger quant for attention tensors.

It's ok if they never release a BF16 model, but it's less ok if they release it, win the benchmarks, then quantise it after a few weeks.

Re: Arena AI Model ELO History

#12

FYI, Elo isn't an acronym - it's a person's name. No need to capitalize it as ELO.

You’re right: https://en.wikipedia.org/wiki/Elo_rating_system

> ELO ratings

Thank you, I just looked at the chart and said to myself: ELO? YOLO!

That Elo ranking is also called chess ranking

Re: Arena AI Model ELO History

#13
The interesting thing I find is how Anthropic has been more consistently improving over time in the last few years, that allows it to catchup and surpass OpenAI and Google. The latter two have pretty much plateau over the last year or so. GPT 5.5 is somehow not moving the needle at all.

I hope to see the other labs can bring back competition soon!

Re: Arena AI Model ELO History

#16
post #13

The interesting thing I find is how Anthropic has been more consistently improving over time in the last few years, that allows it to catchup and surpass OpenAI and Google. The latter two have pretty much plateau over the last year or so. GPT 5.5 is somehow not moving the needle at all. I hope to see the other labs can bring back competition soon!

Gpt 5.5 is quite a big leap, it's a lot better than opus 4.7 for agentic coding

Re: Arena AI Model ELO History

#17

Earlier quoted context omitted.

You’re right: https://en.wikipedia.org/wiki/Elo_rating_system

> ELO ratings Thank you, I just looked at the chart and said to myself: ELO? YOLO! That Elo ranking is also called chess ranking

Élő. Meaning alive (él = it lives, -ő = adjective)

Re: Arena AI Model ELO History

#18
This is great, but personally, I really wish we had an Elo leaderboard specifically for the quality of coding agents.

Honestly, in my opinion, GPT-5.5 Codex doesn't just crush Claude Code 4.7 opus —it's writing code at a level so advanced that I sometimes struggle to even fully comprehend it. Even when navigating fairly massive codebases spanning four different languages and regions (US, China, Korea, and Japan), Codex's performance is simply overwhelming.

How would we even go about properly measuring and benchmarking the Elo for autonomous agents like this?

Re: Arena AI Model ELO History

#19
post #16
post #13

The interesting thing I find is how Anthropic has been more consistently improving over time in the last few years, that allows it to catchup and surpass OpenAI and Google. The latter two have pretty much plateau over the last year or so. GPT 5.5 is somehow not moving the needle at all. I hope to see the other labs can bring back competition soon!

Gpt 5.5 is quite a big leap, it's a lot better than opus 4.7 for agentic coding

Arena only allows very small context sizes, so it's a noisy benchmark for what we care about IRL.

Re: Arena AI Model ELO History

#20
post #18

This is great, but personally, I really wish we had an Elo leaderboard specifically for the quality of coding agents. Honestly, in my opinion, GPT-5.5 Codex doesn't just crush Claude Code 4.7 opus —it's writing code at a level so advanced that I sometimes struggle to even fully comprehend it. Even when navigating fairly massive codebases spanning four different languages and regions (US, China, Korea, and Japan), Cod…

Isn't code that you fail to understand literally a sign that its worse?
Post reply on HN