Before getting too excited, take a look at the intelligence vs cost matrix: https://artificialanalysis.ai/models?intelligence-index-toke...
5.6 Sol (max) being cheaper than all of these is wild, considering how good the output is too
Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
11–20 of 251 posts
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#12Before getting too excited, take a look at the intelligence vs cost matrix: https://artificialanalysis.ai/models?intelligence-index-toke...
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#13On my end, Opus 5 is Haiku level vs. Opus 4.8 (good) and Fable (superb). Gets confused by permission prompts, cannot debug a failing test it caused (Opus 4.8 got it right after, without tens of rounds "thinking").
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#14AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer.
This seems to be a good proxy for param size/density and the ranking breaks down as such: Claude Fable 5 (with fallback), Gemini 3.1 Pro Preview, Claude Opus 5 (Max), Grok 4.6 (high), Gemini 3.6 Flash, GPT 5.6 Sol (Max)
I've thought for a while that Gemini 3.x has 'big model smell'
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#15Earlier quoted context omitted.
5.6 Sol (max) being cheaper than all of these is wild, considering how good the output is too
It shouldn't be surprising OpenAI does have the most compute out of all the major labs. The only reason why Anthropic models are expensive is they are the most in demand models in the world and Anthropic is fighting for compute. The only way to you limit demand for your model is increasing API pricing this is also why Anthropic probably has great margin and probably is profitable compared to OpenAI.
No wonder why Tibo can afford to hit the reset button liberally.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#16On my end, Opus 5 is Haiku level vs. Opus 4.8 (good) and Fable (superb). Gets confused by permission prompts, cannot debug a failing test it caused (Opus 4.8 got it right after, without tens of rounds "thinking").
Are you using Claude Code/CoWork or an API client? I’m curious if it has different training that makes it more effective with specific instructions/ tool calling methods that are only implemented in official harnesses.
I bizarrely had Opus 4.8 this week (in pi.dev within a podman container, using openrouter) start installing various python packages (and uv!) within the environment (not as root) when I asked it to code review some fairly basic Rust .rs files that were generally stand-alone (it did very nicely work out and write some stubs for them to build them and work out how they worked).
It only gave up with the weird Python installing stuff when it discovered one of the Python packages needed Tensorflow.
It seems pretty focused and persistent in continuing its initial approach, and I'm wondering if I need to alter some instructions / initial prompts to rein it in a bit...
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#17Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#18Earlier quoted context omitted.
5.6 Sol (max) being cheaper than all of these is wild, considering how good the output is too
It shouldn't be surprising OpenAI does have the most compute out of all the major labs. The only reason why Anthropic models are expensive is they are the most in demand models in the world and Anthropic is fighting for compute. The only way to you limit demand for your model is increasing API pricing this is also why Anthropic probably has great margin and probably is profitable compared to OpenAI.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#19Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#20At least two models (GPT-5.6, Kimi K3) match its score (~1-2% diff) for half the cost.