Live data from Hacker News

Claude Opus 5

anthropic.com

191–200 of 1001 posts

Re: Claude Opus 5

#191
post #18

Earlier quoted context omitted.

Fable is twice the price.

If Fable gets correct answer quicker, then you might pay less than doing back and forth with Opus, plus you lose more of your own time. I see no reason for using less able models in my workflows. There is this saying, penny wise and pound foolish

less expensive per task might also mean less of your own time

Re: Claude Opus 5

#192
> This means that the model can support defensive cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively.

Annoyingly, this is a concrete argument that open source software may be easier to attack.

Re: Claude Opus 5

#194

That's a crazy arc 3 score. What do people think of this? Are models actually developing fluid intelligence like what the creators claim to be measuring? Is it jus do to training for it? Is the benchmark flawed?

Yes, I think it indicates real progress in fluid intelligence. Clearly these models are making huge strides in usefulness which are well correlated with their ARC-AGI scores. I don't think this is benchmaxxing. These companies are locked in a competition to produce the best software engineer, and falling behind is an existential risk. I doubt they are wasting time benchmaxxing ARC-AGI.

nah they could make educated guess about arc and benchmaxx it too.

Re: Claude Opus 5

#196

Where does this leave Fable? I am confused.

on the API for people who don't want to change models, but I imagine most people will probably switch to their cheaper Opus 5 (cheaper for us and presumably also cheaper for them)

Re: Claude Opus 5

#197
post #26

Earlier quoted context omitted.

Funny that a company selling an AI software developer can't use it to fix their infra. Fixing those issues still requires humans.

Let's be honest - they're also still hiring software devs. AI still requires skilled humans in the loop and that's not going away.

It is going away for non tech companies though.

Imagine you are a company that sells concrete. You have a web dev contractor you use to build and maintain your website. It has tools on it to get delivery quotes and a few internal tools to track orders.

Except now you can just have your sales team also maintain the website with a $20/month Claude subscription.

Re: Claude Opus 5

#198

Wow, 30% on ARC-AGI-3 for $20k total. Huge jump from GPT-5.6's 7.8% at $20k per task. I continue to believe ARC-AGI measures something different and important compared to other benchmarks.

> I continue to believe ARC-AGI measures something different

why is that? its now being benchmaxxed too

Re: Claude Opus 5

#199

Is Fable 5 just Opus 5 with some additional long-context management modifications for extended self-directed work? Or are they actually truly different models?

I suspect they make a big model first. In this case it's Fable. Then they run the shrinker steps to make Sonnet and Opus. Sonnet is smaller, takes less time to make, so it got released first. Opus needed few more weeks to cook. With this iteration they had a delay because when the Mythos was ready they had some sort of "Oh shit" moment and spent half a year adding safety guards to it. Then slowly rolled it out, but g…

But I wonder how they were able to release Sonnet 5 during the period when even people inside Anthropic were legally barred from using Mythos/Fable?

Re: Claude Opus 5

#200
post #183

Wow, 30% on ARC-AGI-3 for $20k total. Huge jump from GPT-5.6's 7.8% at $20k per task. I continue to believe ARC-AGI measures something different and important compared to other benchmarks.

seeing a jump this big is not a great sign for the continuing value of a benchmark

It will continue to be valuable as a cost and speed benchmark long after it is saturated at the high end. And they are already working on ARC-AGI 4 and thinking about going even farther.
Post reply on HN