Karpathy gave his initial impression: https://x.com/karpathy/status/1891720635363254772 The pull quote is: The impression overall I got here is that this is somewhere around (OpenAI) o1-pro capability
The impression seems to be warranted: Grok 3 has directly jumpted to the top of all leaderboard categories in Chatbot Arena: https://lmarena.ai/?leaderboard In math it shares the top spot with o1 and is just a few points behind (well within errors). In creative writing it is basically ex-aequo with the latest ChatGPT 4o and in coding it's actually significantly ahead of everyone else and represents a new SOTA.
Do we have a way to tell if one model is smarter than another at that point?