Chatbot Arena: Benchmarking LLMs in the Wild with Elo Ratings
1–10 of 10 posts
Re: Chatbot Arena: Benchmarking LLMs in the Wild with Elo Ratings
#2You may also learn from 7.33 dota update which uses a new ranking algorithm called Glicko.
Re: Chatbot Arena: Benchmarking LLMs in the Wild with Elo Ratings
#3It's glad to see the old technique is used for new models. You may also learn from 7.33 dota update which uses a new ranking algorithm called Glicko.
Re: Chatbot Arena: Benchmarking LLMs in the Wild with Elo Ratings
#4Re: Chatbot Arena: Benchmarking LLMs in the Wild with Elo Ratings
#5It's glad to see the old technique is used for new models. You may also learn from 7.33 dota update which uses a new ranking algorithm called Glicko.
could you provide any reference? Is it a variant of ELO?
Valve listed some reason for making the change.
Re: Chatbot Arena: Benchmarking LLMs in the Wild with Elo Ratings
#6Surprised to learn StableLM is worse than plain LLaMA. link to their leaderboard: leaderboard.lmsys.org
Re: Chatbot Arena: Benchmarking LLMs in the Wild with Elo Ratings
#7Re: Chatbot Arena: Benchmarking LLMs in the Wild with Elo Ratings
#8Re: Chatbot Arena: Benchmarking LLMs in the Wild with Elo Ratings
#9Re: Chatbot Arena: Benchmarking LLMs in the Wild with Elo Ratings
#10It's glad to see the old technique is used for new models. You may also learn from 7.33 dota update which uses a new ranking algorithm called Glicko.
could you provide any reference? Is it a variant of ELO?
One is a rating system named after the creator, Arpad Elo. See https://en.wikipedia.org/wiki/Elo_rating_system
The other is a rock band that was formed in 1970. See https://en.wikipedia.org/wiki/Electric_Light_Orchestra