Arena AI Model ELO History
41–50 of 63 posts
Re: Arena AI Model ELO History
#42The Elo rating system measures relative performance to the other models. As the other models improve or rather newer better models enter the list, the Elo score of a given existing model will tend to decrease even though there might be no changes whatsoever to the model or its system prompt. You can't use Elo scores to measure decay of a models performance in absolute terms. For that you need a fixed harness running…
Re: Arena AI Model ELO History
#43Is this slop? It has wildly aggressive language that agrees with a subset of pop sentiment, re: models being “nerfed”. It promises to reveal this nerfing. Then, it goes on to…provide an innocuous mapping of LM Arena scores that always go up?
Re: Arena AI Model ELO History
#44Earlier quoted context omitted.
No, that is not how ELO scores work.
As far as I understand, this is exactly how ELO scores work. If a more capable show up and starts beating all the other models, it literally takes ELO points from everyone else. https://en.wikipedia.org/wiki/Elo_rating_system
Re: Arena AI Model ELO History
#45> the slow performance decays the decays are just more capable other models entering the population, making all prior models lose more frequently
No, that is not how ELO scores work.
Re: Arena AI Model ELO History
#46The Elo rating system measures relative performance to the other models. As the other models improve or rather newer better models enter the list, the Elo score of a given existing model will tend to decrease even though there might be no changes whatsoever to the model or its system prompt. You can't use Elo scores to measure decay of a models performance in absolute terms. For that you need a fixed harness running…
Is that strictly true? ELO rankings do also inflate over time (looking at you, Chess GMs)
Re: Arena AI Model ELO History
#47Seems like Chinese labs are the only ones that are trustworthy (at least when it gets to this specific issue). This feels so ironic haha
Re: Arena AI Model ELO History
#48Earlier quoted context omitted.
There's almost 0% chance that OpenAI doesn't quantize the model right off the bat. I am willing to bet large amounts of money that OpenAI would never release a model served as fully BF16 in the year of our lord 2026. That would be insane operationally. They're almost certainly doing QAT to FP4 for FFN, and a similar or slightly larger quant for attention tensors.
It's ok if they never release a BF16 model, but it's less ok if they release it, win the benchmarks, then quantise it after a few weeks.
Re: Arena AI Model ELO History
#49It seems to be a USA only thing, Chinese models and Mistral don't show any downward trend.
Re: Arena AI Model ELO History
#50Earlier quoted context omitted.
No, that is not how ELO scores work.
As far as I understand, this is exactly how ELO scores work. If a more capable show up and starts beating all the other models, it literally takes ELO points from everyone else. https://en.wikipedia.org/wiki/Elo_rating_system
If a more capable show up and starts
beating all the other models
There is an instance of this in the chart. In 2025-06-24 when Gemini-2.5-pro shows up. As you can see, the ELO of the others do not drop.