Claude Sonnet 5 – benchmark results
artificialanalysis.ai
Claude Sonnet 5 – benchmark results
1–10 of 20 posts
Re: Claude Sonnet 5 – benchmark results
#2Half of the data is missing and the rest is inconsistent between different graphs and sections. Is the benchmark having Sonnet 5 generate the page and seeing how many hallucinations it has?
Re: Claude Sonnet 5 – benchmark results
#3Seems like the model is incredibly inefficient at max reasoning, and even at high/xhigh it uses far more tokens than other models, including Gemini 3.5 Flash, GLM 5.2 and so on. GPT 5.5's efficiency in tokens is still unmatched.
See also: https://cursor.com/cursorbench
Re: Claude Sonnet 5 – benchmark results
#4Seems like the model is incredibly inefficient at max reasoning, and even at high/xhigh it uses far more tokens than other models, including Gemini 3.5 Flash, GLM 5.2 and so on. GPT 5.5's efficiency in tokens is still unmatched. See also: https://cursor.com/cursorbench
Same with opus nothing above medium has a reasonable improvement for the tokens spent.
Re: Claude Sonnet 5 – benchmark results
#5Yet another mediocre model. Mostly irrelevant among open weights alternatives. Fable wen.
Re: Claude Sonnet 5 – benchmark results
#6I'm so sick of Anthropics usage caps and how their model devours tokens.
Re: Claude Sonnet 5 – benchmark results
#7Using Fable, pretty much every request hit some gate they had for no discernible reason. These provider-level rejections should be incorporated into benchmarks as 0s on the tasks since that's the experience you'll actually get using the model.
Re: Claude Sonnet 5 – benchmark results
#8I'm so sick of Anthropics usage caps and how their model devours tokens.
It starts with NVIDIA artificially and slowly releasing its tech. If the GPUs were cheaper, we would have better models by many other companies, and competition would take care of these greedy tactics.
Re: Claude Sonnet 5 – benchmark results
#9Cost per task is shockingly high. More expensive than Opus 4.8, second in place to Fable.
Cost per task data is only available for max effort though, might just be very inefficient at that effort level.
Re: Claude Sonnet 5 – benchmark results
#10I feel like they repackaged Opus, slightly nerfed it, and reduced price per token.
A release just to have a headline while Fable situation is getting resolved.