Live data from Hacker News

Claude Opus 5

anthropic.com

221–230 of 1001 posts

Re: Claude Opus 5

#221
The truth for me at least is that these models became "good enough" around Opus 4.6. I feel like further capability improvements, "step changes" like we saw with agentic coding, aren't necessarily going to come from the model. I think the next crown goes to whoever can figure out the right scaffolding so that these models can be inserted into your organization.

Maybe I'm wrong and Opus 5 is a real unlock?

Re: Claude Opus 5

#223
post #164

Their communication is confusing. They say "Opus 5 is not more capable overall than Fable 5", but their blog post proceeds to list how much better Opus 5 is than Fable 5 on __most__ benchmarks listed. Then system card goes on to "Its AI R&D capabilities are comparable to those of Claude Mythos 5", which is supposed to be fable minus restrictions.

Easy enough to explain: they're benchmaxxing. Fable is intelligent but not benchmaxxed. Opus is less intelligent but benchmaxxed.

Re: Claude Opus 5

#224
post #104

I'm not sure what to make of this graph[0]. It shows medium as the most effective thinking mode by far for frontier code. It's the only case that I saw going through the system card where more reasoning effort meaningfully negatively impacted the resulting eval. I know sometimes max efforts show a small dip, but this is substantial. I wonder why in the world that is? [0] https://imgur.com/a/Nv8V7Ry

apparently it got docked points for editing files out of scope

Re: Claude Opus 5

#225
post #23

Great that there's a new model but they could fix their existing infra. We're considering dropping our Claude Team sub cause it's unusable recently. Constant bugs, dropped sessions, issues switching models, http errors. It's becoming ridiculous

> Constant bugs, dropped sessions, issues switching models, http errors. It's becoming ridiculous

And memory leaks.

Re: Claude Opus 5

#226
Can someone help me understand something? I thought Fable was such a miraculous leap forward in capability. But now it seems Opus is basically on par with it, and in some cases (computer use) far exceeds it.

Re: Claude Opus 5

#227
GPT 5.6 Sol is the first model I've used where I can trust it to add 100-500 lines of code maintainably.

It's great with Codex.

I still find that LLMs tend to not know how to compose larger ideas but on the scale of small ideas or short form well defined tasks like small scale debugging/performance engineering it's safe to say that they are now superhuman.

Re: Claude Opus 5

#228
post #155

I think the most important thing here is not absolute performance. It's that organizations now have access to a Fable-ish model without Fable's 30-day data retention requirement[0]. > "Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access."[1] On the Opus model release page, the reason why Fable doesn't have an ARC-AGI score is because of that retention policy[2]. 0: h…

Also the cost per task. It appears to be significantly cheaper, cheaper than sonnet!

The numbers from Anthropic seem heavily cherry-picked, Artificial Analysis has Opus 5 at 1.25x the cost of Sonnet and 2x the cost of GPT 5.6 and K3.

https://artificialanalysis.ai/?cost=cost-per-task

Re: Claude Opus 5

#229
post #104

I'm not sure what to make of this graph[0]. It shows medium as the most effective thinking mode by far for frontier code. It's the only case that I saw going through the system card where more reasoning effort meaningfully negatively impacted the resulting eval. I know sometimes max efforts show a small dip, but this is substantial. I wonder why in the world that is? [0] https://imgur.com/a/Nv8V7Ry

I agree, very odd they did not comment on any theories for the degradation here. Dip and then rebound at max effort is pretty interesting too. Overthinking is bad, but you can overthink so much it starts to be better again?

Re: Claude Opus 5

#230
post #104

I'm not sure what to make of this graph[0]. It shows medium as the most effective thinking mode by far for frontier code. It's the only case that I saw going through the system card where more reasoning effort meaningfully negatively impacted the resulting eval. I know sometimes max efforts show a small dip, but this is substantial. I wonder why in the world that is? [0] https://imgur.com/a/Nv8V7Ry

[deleted]
Post reply on HN