The LMSys Overall leaderboard https://chat.lmsys.org/?leaderboard > can tell us a bit more about how these models will perform in real life, rather than in a benchmark context. By comparing the ELO score against the MMLU benchmark scores, we can see models which outperform / underperform based on their benchmark scores relative to other models. A low score here indicates that the model is more optimized for the bench…
These days, lmsys elo is the only thing I trust. The other benchmark scores mean nothing to me at this point
For my use of the chat interface, I don't think lmsys is very useful. lmsys mainly evaluates relatively simple, low token count questions. Most (if not all) are single prompts, not conversations. The small models do well in this context. If that is what you are looking for, great. However, it does not test longer conversations with high token counts.
Just saying that all benchmarks, including lmsys, have issues and are focused on specific use cases.