Live data from Hacker News

Claude Opus 5

anthropic.com

521–530 of 1001 posts

Re: Claude Opus 5

#521

I think the most important thing here is not absolute performance. It's that organizations now have access to a Fable-ish model without Fable's 30-day data retention requirement[0]. > "Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access."[1] On the Opus model release page, the reason why Fable doesn't have an ARC-AGI score is because of that retention policy[2]. 0: h…

[flagged]

Re: Claude Opus 5

#522

What's the point of 150 pages description of a model that's going to be replaced in a couple months? Who even reads this? I know it's cheap to generate text with LLMs, but this is just noise at this point.

These cards are very important to knowing what the companies are tracking regarding safety.

Not just in what the models can or might want to do, but how they treat the operators they interact with.

If you look carefully, this card shows the addition of a new benchmark for "condescension" as a character trait.

I think a lot of people would like to see a comparable system card for the unannounced model that escaped openai last week.

Re: Claude Opus 5

#523

That's a crazy arc 3 score. What do people think of this? Are models actually developing fluid intelligence like what the creators claim to be measuring? Is it jus do to training for it? Is the benchmark flawed?

It passed the first two puzzles, which are incredibly simple but the bench doesn't explain what the goal is. Any model with a knowledge cut-off after the introduction of ARC-AGI-3 could probably pass the first two puzzles just by knowing what the goal is.

Re: Claude Opus 5

#524
post #164

Their communication is confusing. They say "Opus 5 is not more capable overall than Fable 5", but their blog post proceeds to list how much better Opus 5 is than Fable 5 on __most__ benchmarks listed. Then system card goes on to "Its AI R&D capabilities are comparable to those of Claude Mythos 5", which is supposed to be fable minus restrictions.

Easy enough to explain: they're benchmaxxing. Fable is intelligent but not benchmaxxed. Opus is less intelligent but benchmaxxed.

That's a plausible explanation but I'm not seeing evidence for it.

I have a personal benchmark suite of 14 real, non-public tasks. Opus 5 and Fable tied on 10, Opus won on 3, and Fable won on 1. It's a really strong model.

Re: Claude Opus 5

#526
post #505

Seems to me the purpose of all these releases, credits, pricing changes, harness changes, unpredictable token usages for the same task, etc. is to keep customers completely befuddled so that it's impossible to compare AI products. It's like hiring a consultant who sends invoices every month that aren't related to hours worked or project progress, but are whatever the consultant feels like billing, and you're expected…

It's a new and improved version of an existing model? I don't think it's intentionally befuddling.

Re: Claude Opus 5

#527

Earlier quoted context omitted.

Switch to K3 and you won’t look back, I promise!! I got so fed up with Claude and finally bit the bullet to switch and it’s amazing

I really want to, but I don't have the cluster at home, and I don't use token-based billing except at DeepSeek prices.

Real… I’ve been using Deepseek v4 pro max as my main and then k3 in web (more usage credits) for automating what my deepseek agents do

Re: Claude Opus 5

#528
This is cool, but I wish we could stop building landing pages to assess the intelligence of these models. There is much more to them than that. There are infinite number of complicated things that require a crap-ton of intelligence (biological or digital). The most fascinating of these for me these days is large scale migrations. Projects that are so ginormous and risky that many teams have either given up on them, or don't get funding. But with models like Opus, those projects are now within reach. What's MORE fascinating is that leadership is now asking if we can use opus models to get the refactor/migration done. This is the opposite of what has been happening for decades. Its so hard to make a convincing and affordable business case for large scale refactoring or migrations.

Re: Claude Opus 5

#529
post #446

I compared the writing style of Opus 5 vs Fable 5, and Opus 5 continues many of the "Claude-isms" of its 4.8 predecessor in a way that Fable broke away from. Opus 5 still uses "carry the argument", "worth stating plainly", ", and the trap", "The X matters more", the use of "move" We need an "annoying English" benchmark. - Fable 5 Max: https://gist.github.com/deet/3d97f854b48eac6658d642fa18bb24d... - Opus 5 Max: https…

Fable might be using those phrases less, but its writing is still terrible and exhausting to read.

Agreed. I’ve interestingly found 5.6 sol to produce much better writing, and it can generally cut to the point much more effectively.

Re: Claude Opus 5

#530
post #446

I compared the writing style of Opus 5 vs Fable 5, and Opus 5 continues many of the "Claude-isms" of its 4.8 predecessor in a way that Fable broke away from. Opus 5 still uses "carry the argument", "worth stating plainly", ", and the trap", "The X matters more", the use of "move" We need an "annoying English" benchmark. - Fable 5 Max: https://gist.github.com/deet/3d97f854b48eac6658d642fa18bb24d... - Opus 5 Max: https…

I've found Opus 4.8 and Fable 5 both difficult to learn from purely because of how annoying their writing style is. I'm finding GPT 5.6 Sol to be much better for this.
Post reply on HN