Arena AI Model ELO History
51–60 of 63 posts
Re: Arena AI Model ELO History
#52FYI, Elo isn't an acronym - it's a person's name. No need to capitalize it as ELO.
Re: Arena AI Model ELO History
#53The logic for which models stay active when you click on a group of them is extremely not working
Re: Arena AI Model ELO History
#54The Elo rating system measures relative performance to the other models. As the other models improve or rather newer better models enter the list, the Elo score of a given existing model will tend to decrease even though there might be no changes whatsoever to the model or its system prompt. You can't use Elo scores to measure decay of a models performance in absolute terms. For that you need a fixed harness running…
Re: Arena AI Model ELO History
#55Is this slop? It has wildly aggressive language that agrees with a subset of pop sentiment, re: models being “nerfed”. It promises to reveal this nerfing. Then, it goes on to…provide an innocuous mapping of LM Arena scores that always go up?
Re: Arena AI Model ELO History
#56> the slow performance decays the decays are just more capable other models entering the population, making all prior models lose more frequently
No, that is not how ELO scores work.
Re: Arena AI Model ELO History
#57FYI, Elo isn't an acronym - it's a person's name. No need to capitalize it as ELO.
Re: Arena AI Model ELO History
#58For what it's worth, I work at OpenAI and I can guarantee you that we don't switch to heavily quantized models or otherwise nerf them when we're under high load. It's true that the product experience can change over time - we're frequently tweaking ChatGPT & Codex with the intention of making them better - but we don't pull any nefarious time-of-day shenanigans or similar. You should get what you pay for.
> we don't switch to heavily quantized models That sounded like a press bulletin, so just to let you clarify yourself: Does that mean you may switch to lightly quantized models?
(I used the adjective heavily because that’s what the original post said. I have no intention of making misleading but technically true statements.)
Re: Arena AI Model ELO History
#59Earlier quoted context omitted.
It's ok if they never release a BF16 model, but it's less ok if they release it, win the benchmarks, then quantise it after a few weeks.
that is for sure what everyone does. also they train on evals with the datasets that they would be bench against.
(The loose version of this that’s true is that there may exist eval data contamination in pretraining. This is a hard problem to fully solve.)
Re: Arena AI Model ELO History
#60Earlier quoted context omitted.
that is for sure what everyone does. also they train on evals with the datasets that they would be bench against.
What do you mean by this? We don’t train on evals, and if we did I’d quit on the spot. (The loose version of this that’s true is that there may exist eval data contamination in pretraining. This is a hard problem to fully solve.)