Live data from Hacker News

Claude Opus 5

anthropic.com

131–140 of 1001 posts

Re: Claude Opus 5

#131

Rather interesting that this makes sonnet 5 look even worse! There is no reason to use sonnet over opus with low or no reasoning at all.

My understanding is that Opus should be used for planning, macro-level conversations and Sonnet for execution.

So, for coding, for example: Opus for solution design and architectural blueprint and then Sonnet for actual implementation.

Works out cheaper with minimal loss of quality.

At least that's my personal understanding and anecdotal experience.

Re: Claude Opus 5

#132

Edit: It was pointed out to me that Opus 4.8 got "21%" for successfully fully completing ~1-in-5 tasks, but also got "55.7%" for obtaining significant partial credit on some of the ~4-in-5 tasks it could not fully complete. --------------- Why does Anthropic say here that Opus 4.8 scored 55.7% on OSWorld 2.0 benchmark, but the paper published by the authors of OSWorld 2.0 say they achieved a benchmark of ~21% with Op…

You're comparing the "score percentage" (e.g. out of the total number of partial points available, how many did the agent achieve) to the "completion percentage" (how many tasks does the model score 100% on). The paper says "Claude Opus 4.8 with maximum thinking and batched tool calls scores best but still completes only 20.6% of tasks at a 54.8% partial score", which is ~the same number that Anthropic reports here (55.7 vs 54.8).

That is—the agent scored 100% on 20% of tasks, but on average it got 54% of the "score" awarded in the exam. One number reflects partial progress, the other one doesn't. The authors of the benchmark prefer you to look at the lower number (because they want to show their benchmark as capturing useful gaps in capabilities and with a lot of room for improvement), the authors of the models want you to look at the higher number (because they want you to think of their models as capable)

Re: Claude Opus 5

#134
post #26
post #23

Great that there's a new model but they could fix their existing infra. We're considering dropping our Claude Team sub cause it's unusable recently. Constant bugs, dropped sessions, issues switching models, http errors. It's becoming ridiculous

Funny that a company selling an AI software developer can't use it to fix their infra. Fixing those issues still requires humans.

This is just like any extreme engineering domain. I am okay with occasional delays in Flights, as long as it takes me from X to Y in 10hrs vs months.

Re: Claude Opus 5

#135
I think the most important thing here is not absolute performance. It's that organizations now have access to a Fable-ish model without Fable's 30-day data retention requirement[0].

> "Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access."[1]

On the Opus model release page, the reason why Fable doesn't have an ARC-AGI score is because of that retention policy[2].

0: https://support.claude.com/en/articles/15425996-data-retenti...

1: https://www.anthropic.com/news/claude-opus-5

2: https://xcancel.com/arcprize/status/2064399134099153344

Re: Claude Opus 5

#136

Earlier quoted context omitted.

Is that true? Sol responses are also longer than prior models.

Anecdata: my workflow has been working on the same personal projects for months now with Codex. I cannot anymore finish my daily/weekly code with 4.8 anymore. I was dividing my work between Codex and DeepSeek. Now I barely use DeepSeek, or never because Codex quota is enough after Sol

I hit Codex limits (20x account, never using /fast) on Sol Medium in about 2.5 days

Re: Claude Opus 5

#137

Edit: It was pointed out to me that Opus 4.8 got "21%" for successfully fully completing ~1-in-5 tasks, but also got "55.7%" for obtaining significant partial credit on some of the ~4-in-5 tasks it could not fully complete. --------------- Why does Anthropic say here that Opus 4.8 scored 55.7% on OSWorld 2.0 benchmark, but the paper published by the authors of OSWorld 2.0 say they achieved a benchmark of ~21% with Op…

[deleted]

Re: Claude Opus 5

#138

That's a crazy arc 3 score. What do people think of this? Are models actually developing fluid intelligence like what the creators claim to be measuring? Is it jus do to training for it? Is the benchmark flawed?

It’s still “only” at 30%, and “fluid intelligence” isn’t very well-defined. The models are getting more capable, but what that means in absolute terms is anyone’s guess, because we don’t have a thorough understanding on what exactly constitutes human intelligence.

I’d say the proof is in the pudding, that is, in real-world applications. We are still seeing important limitations in LLMs.

Re: Claude Opus 5

#139
post #124

Edit: It was pointed out to me that Opus 4.8 got "21%" for successfully fully completing ~1-in-5 tasks, but also got "55.7%" for obtaining significant partial credit on some of the ~4-in-5 tasks it could not fully complete. --------------- Why does Anthropic say here that Opus 4.8 scored 55.7% on OSWorld 2.0 benchmark, but the paper published by the authors of OSWorld 2.0 say they achieved a benchmark of ~21% with Op…

I think some variance is to be expected since LLMs are typically non-deterministic, however that's a huge difference that I think warrants further explanation.

[deleted]

Re: Claude Opus 5

#140

Edit: It was pointed out to me that Opus 4.8 got "21%" for successfully fully completing ~1-in-5 tasks, but also got "55.7%" for obtaining significant partial credit on some of the ~4-in-5 tasks it could not fully complete. --------------- Why does Anthropic say here that Opus 4.8 scored 55.7% on OSWorld 2.0 benchmark, but the paper published by the authors of OSWorld 2.0 say they achieved a benchmark of ~21% with Op…

[deleted]
Post reply on HN