Grok3 is first model to surpass 1400 on the Chat Arena benchmark
1–4 of 4 posts
Re: Grok3 is first model to surpass 1400 on the Chat Arena benchmark
#2It doesn't feel groundbreaking to me yet — it doesn't feel consistently better than the rest of the frontier models — but it's definitely a frontier model. Congratulations to the xAI team for getting to the frontier so quickly.
Re: Grok3 is first model to surpass 1400 on the Chat Arena benchmark
#3In general the Chatbot Arena leaderboards aren't considered super-reliable benchmarks anymore, especially since they allow model makers to effectively hill-climb on them by pre-releasing models. That being said, Grok 3 has also done quite well on many standard benchmarks; my personal vibecheck of asking some tricky math problems and various Kubernetes architecture questions place it around DeepSeek V3, for me. The th…
Re: Grok3 is first model to surpass 1400 on the Chat Arena benchmark
#4In general the Chatbot Arena leaderboards aren't considered super-reliable benchmarks anymore, especially since they allow model makers to effectively hill-climb on them by pre-releasing models. That being said, Grok 3 has also done quite well on many standard benchmarks; my personal vibecheck of asking some tricky math problems and various Kubernetes architecture questions place it around DeepSeek V3, for me. The th…
Just checking, you were able to try it because you pay right?