Something is odd with this model, their blog posts shows REALLY good results, but in most other third-party benchmarks, people realize it's not really SOTA, even bellow Kimi K2.6 and GLM-5/5.1 In my tests too[0], it doesn't reach top 10. One issue, which they also mentioned in their post, is that they can't really serve well the model at the moment, so V4-Pro is heavily rate-limited and gives a lot of timeout errors…
Hmm, the Flash performs significantly better than Pro in the benchmark? That's very strange; could rate limiting cause that?
I expect once the API issues are fixed, for v4-pro to be around the same level as GLM-5.