The naming system is so confusing. Is Opus better than Sonnet? Where does Haiku fit in? How can you tell from the name? I can't keep track of all these names or make guesses from the names. Suggestion for a better naming system: use the words "Pro", "Plus", etc.: Claude 5 Pro, Claude 5 Standard, Claude 5 Fast, Claude 5 Mini.
Opus is better than Sonnet -- an Opus is longer than a Sonnet
An opus is just short for "magnum opus", and it's a different type of label than a sonnet. A sonnet is a very particular kind of poem, while "opus" basically just means an important work. It can be short or long, and has no format requirements like sonnet (or haiku does).
I compared the writing style of Opus 5 vs Fable 5, and Opus 5 continues many of the "Claude-isms" of its 4.8 predecessor in a way that Fable broke away from. Opus 5 still uses "carry the argument", "worth stating plainly", ", and the trap", "The X matters more", the use of "move" We need an "annoying English" benchmark. - Fable 5 Max: https://gist.github.com/deet/3d97f854b48eac6658d642fa18bb24d... - Opus 5 Max: https…
I think a signature Claude style of writing is good since it makes it that much harder to pass off Claude written text as human.
Here's another test of a cyberpunk ramen shop website. One thing I've found LLMs have a lot of difficulty with is angular cuts / elements that aren't easily representable with CSS. Cyberpunk aesthetics are generally a great test of that, since they have a lot of microglyphs / window decoration. Design source of truth: https://image.non.io/9d5fed20-b476-49d3-841b-37eb553fb88e.we... Opus 5 build: https://html.non.io/ne…
God damn, we are living in the future. I love this so much. Designs like this would never have seen the light of day in the cellphone incrementalism / corporate memphis era of tech. Now people can be weird and awesome again. This is 1980's cyberpunk / late-90's Matrix / early-00's sci-fi UI. Great ideas that died to frutiger aero (which isn't a bad design aesthetic) and flat design (which is). This is fun and it's go…
Unrelated question, what’s your favorite flavor of Kool Aid?
Is it me or these have gotten very boring. We have 5 more points on xyzbench or whatever .
It's you. The benchmarks don't matter much. We have very little hands on experience with this thing yet. Give it a few days, and be cranky then. Right now, it seems it is getting close to Fable level while being 2x cheaper. That's not boring.
This is part of why I switched to Grok 4.5 I don't need more powerful models, I need one that responds fast enough that my attention doesn't wander to other tasks. Grok 4.5 is so fast I can just use it in-band without swapping to other tasks. Slower than Opus 4.8, which was already miserably slow, is indeed a step in the wrong direction.
for you, I am in general not constrained by the speed of models since I parallelize. For me autonomy and accuracy are paramount above all else.
I find that my orchestrating agent still needs to be fast to properly coordinate a fleet of slower models.
During post-training of opus 5, the last few days, opus was a real wreck. I had to swap in gpt 5.6 sol for my orchestrator and enable fast mode (1.5x speed) in order for it to keep up with work and communications from a handful of mostly 5.6 sol agents.
Also because interacting with a slow orchestrator is no fun, even when plenty of work is getting done in parallel in the background.
Also the cost per task. It appears to be significantly cheaper, cheaper than sonnet!
Also the cost per task. https://www.vals.ai/benchmarks/vals_index !!! Vals !!! Vals Index Opus 4.8 > 5.0 goes from $2.90 to $8.54, for 4% gain ... That is a massive cost increase. Sure, 20% cheaper then Fable, but that is a 3x price increase compared to Opus 4.8 in that test. https://artificialanalysis.ai/models/claude-opus-5 https://artificialanalysis.ai/models/claude-opus-5#price-cos... !!! artificial analysis !! C…
Your numbers are for “max”. Opus 5.0 “max” is $2.03. Opus 5.0 “high” (competitive with Claude 4.8 “max” on that index) is $1.06, less than the $1.80 you are quoting for 4.8 max.
That the most expensive variant is expensive doesn’t really tell us much.
Wait, 30% on ARC-AGI-3! I definitely didn't expect that jump so soon. Are there any rumors of what they are changing in architecture that is leading to this?
30% on ARC-AGI-3 is the first two puzzles. It cost $20000 in tokens to do that. That is a terrible result that doesn't imply anything.
Is it some Claude Team/Enterprise only problematic? I'm using two 20x Max accounts almost non-stop (Fable/Opus) for 1.5 years at this point, zero issues with both client and infra sides (from US and in travels). When I'm reading such messages it feels like either I'm lucky or it's a part of some campaign.
I’m on the biggest max plan. It is riddled with annoying bugs for me, only been using it for a little over a month. Settings screen flashes randomly. But most annoyingly: sometimes when forking chats or sometimes for no reason, the UI just straight up eats my previous messages. The model is still aware of them and can recount them if I ask but the visible history is gone. And that’s not even all of them. Fable 5 is j…
I've been on one to two of these plans for eight months or so. There used to be a lot of issues with CC's terminal but at least in iterm2 they have largely been sorted.
And to the guardrails of Fable: https://x.com/cheatyyyy/status/2080693704290140330
That tweet says: > Opus 5 can silently fallback to Opus 4.8 (without any notice) on the serverside if you hit a guardrail But https://support.claude.com/en/articles/16049681-why-claude-s... says (emphasis mine): > These checks cause Claude to _visibly_ fallback from Opus 5 to Opus 4.8 [...] You'll see a notice explaining that the model switched, and the response will be labeled with the model that answered. So who is…
So, announcing the fallback is better than doing it silently, but the fact that Fable falls back frequently for the kind of work I do (a lot of security oriented stuff lately, but it falls back on seemingly random stuff, sometimes, too), means I reach for it less. Getting interrupted mid-task makes it much less valuable. If I have any suspicion I'm going to hit the guardrails, I'll use something else.