I’ve got my stuff rigged to hit mixtral-8x7, and dolphin locally, and 3.5-turbo, and the 4-series preview all with easy comparison in emacs and stuff, and in fairness the 4.5-preview is starting to show some edge on 8x7 that had been a toss-up even two weeks ago. I’m still on the mistral-medium waiting list. Until I realized Perplexity will give you a decent amount of Mistral Medium for free through their partnership…
On what metrics? LMSys shows it does well but 4-Turbo is still leading the field by a wide margin.
I am using 8x-7b internally for a lot of things and Mistral-7b fine-tunes for other specific applications. They're both excellent. But neither can touch GPT-4-turbo (preview) for wide-ranging needs or the strongest reasoning requirements.
https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...
EDIT: Neither does mistral-medium, which I didn't discuss, but is in the leaderboard link.