30% on ARC-AGI-3
Claude Opus 5
181–190 of 1001 posts
Re: Claude Opus 5
#182> Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source-code vulnerability discovery at all access levels. This means that the model can support defensive cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively. Okay so it’s worse than Opus 4.8 for my purposes I guess?
What are your purposes?
Re: Claude Opus 5
#183Wow, 30% on ARC-AGI-3 for $20k total. Huge jump from GPT-5.6's 7.8% at $20k per task. I continue to believe ARC-AGI measures something different and important compared to other benchmarks.
Re: Claude Opus 5
#184Their communication is confusing. They say "Opus 5 is not more capable overall than Fable 5", but their blog post proceeds to list how much better Opus 5 is than Fable 5 on __most__ benchmarks listed. Then system card goes on to "Its AI R&D capabilities are comparable to those of Claude Mythos 5", which is supposed to be fable minus restrictions.
Re: Claude Opus 5
#185Earlier quoted context omitted.
Funny that a company selling an AI software developer can't use it to fix their infra. Fixing those issues still requires humans.
How many companies at the size of Anthropic can serve the amount of traffic and manage the amount of compute they have?
Re: Claude Opus 5
#186Earlier quoted context omitted.
Go read the safeguards section in the report and you will realize why that is. These models are heavily as safeguarded and that was the initial reason why they said they couldn't and haven't released Mythos because that model is the one without the safeguards. OpenAI is did the same thing when they announced a model without safeguards broken into HuggingFace servers.
Yes, this makes a lot of sense, but it’s just very amusing to see. 2 months ago, the world was about to end, now not so much.
Re: Claude Opus 5
#187Is Fable 5 just Opus 5 with some additional long-context management modifications for extended self-directed work? Or are they actually truly different models?
Re: Claude Opus 5
#188i guess the next stuff will be tool use for the rest of what cad does in assemblies and simulation?
itd be fun to try to set up a 3d printer as part of a feedback loop, and see what a model can build.
the automated test harness for physical stuff seems a bit beyond reach still
Re: Claude Opus 5
#189Re: Claude Opus 5
#190"I don't have a reliable way to read that number, so I'd be guessing if I gave you one — and this is exactly the kind of question where a confident guess is worse than none.
What I can tell you is what I actually observe:"
I really like this update - gave me a clear sense of the facts but didn't give me a guess just for the sake of guessing.
One oddity is that it appears to only have a 200K context window right now via CC. Hopefully the 1M version will appear soon!