Live data from Hacker News

Claude Opus 5

anthropic.com

431–440 of 1001 posts

Re: Claude Opus 5

#431
post #164

Their communication is confusing. They say "Opus 5 is not more capable overall than Fable 5", but their blog post proceeds to list how much better Opus 5 is than Fable 5 on __most__ benchmarks listed. Then system card goes on to "Its AI R&D capabilities are comparable to those of Claude Mythos 5", which is supposed to be fable minus restrictions.

Easy enough to explain: they're benchmaxxing. Fable is intelligent but not benchmaxxed. Opus is less intelligent but benchmaxxed.

Honestly that's the simplest explanation and thus likely the correct one.

Re: Claude Opus 5

#432

I found opus 4.8 too agreeable and too wordy(as opposed to codex) and too agreeable. If you are reading documents generating by it was too much. TBH. Fable did a bit better on this. Anyone seen a marked difference with opus 5 on this?

Opus yes, it likes to explain steps and reread files.

Fable is not better, it says zero information between steps and then output a summary. A perfect “send - done”.

Re: Claude Opus 5

#433
post #182

Earlier quoted context omitted.

What are your purposes?

Reversing for the most part, though lately I’ve been doing some code obfuscation/binary rewriting stuff. Fable will switch to Opus instantly on these and I’m unsure how this will perform. I suppose the only way to find out is to test.

Try GPT 5.6 Sol if you haven't yet.

I recently created a patch for Riftborne via static IL patching and Fable 5 outright kept refusing to do it, no issue whatsoever with GPT 5.6 Sol lol.

Re: Claude Opus 5

#434

As a coder, I’ve had no desire to use Fable. In fact I switched from Opus models to sonnet 5 and haven’t noticed any drop in quality on large repos. It seems the gap at the top is very small and not hugely noticeable for backed/frontend. Has anyone else had this experience?

I use Opus for specs and planning, Sonnet for code generation.

Re: Claude Opus 5

#435
post #46

From the prompting guide https://platform.claude.com/docs/en/build-with-claude/prompt... >: > Claude Opus 5's default user-facing responses run longer than prior Opus models'. The benchmarks do show Opus 5 as slightly more expensive than 4.8, although the scores are much higher. This still feels like a step in the wrong direction, though, especially with OpenAI making so much progress with the efficiency of their mod…

This is part of why I switched to Grok 4.5 I don't need more powerful models, I need one that responds fast enough that my attention doesn't wander to other tasks. Grok 4.5 is so fast I can just use it in-band without swapping to other tasks. Slower than Opus 4.8, which was already miserably slow, is indeed a step in the wrong direction.

for you, I am in general not constrained by the speed of models since I parallelize. For me autonomy and accuracy are paramount above all else.

Re: Claude Opus 5

#436
post #104

I'm not sure what to make of this graph[0]. It shows medium as the most effective thinking mode by far for frontier code. It's the only case that I saw going through the system card where more reasoning effort meaningfully negatively impacted the resulting eval. I know sometimes max efforts show a small dip, but this is substantial. I wonder why in the world that is? [0] https://imgur.com/a/Nv8V7Ry

apparently it got docked points for editing files out of scope

This must not be weighted very heavily on the benchmark because if it was, Opus would bomb every test (half kidding)

Re: Claude Opus 5

#439

I have a side project that I always run a simple security analysis prompt on in CC, at each model release. Obviously, Fable 5 would downgrade to Opus 4.8 on any such request. Nothing since Opus 4.6 has found anything interesting. Just ran it using Opus 5, and it found a genuine issue that I verified. Neato!

Do you have a skill for that or do you (or anyone else here) just prompt with "try to find security issues"?

Re: Claude Opus 5

#440

Earlier quoted context omitted.

Easy enough to explain: they're benchmaxxing. Fable is intelligent but not benchmaxxed. Opus is less intelligent but benchmaxxed.

Yep , same with 5.6. Fable is still the best.

But still nerfed compared to the initial release.
Post reply on HN