It ranks between Mistral Small and Mistral Medium on my NYT Connections benchmark and is indeed better than Command R Plus and Qwen 1.5 Chat 72B, which were the top two open weights models. Grok 1.0 is not an instruct model, so it cannot be compared fairly.
Can you share the details about the benchmark?
Most other benchmarks don't clearly show the difference between the top models and the rest. This may be because they are older and have been over-optimized or perhaps because they are just easier.