Live data from Hacker News

Claude Opus 5

anthropic.com

181–190 of 1001 posts

Re: Claude Opus 5

#182
post #38

> Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source-code vulnerability discovery at all access levels. This means that the model can support defensive cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively. Okay so it’s worse than Opus 4.8 for my purposes I guess?

What are your purposes?

Reversing for the most part, though lately I’ve been doing some code obfuscation/binary rewriting stuff. Fable will switch to Opus instantly on these and I’m unsure how this will perform. I suppose the only way to find out is to test.

Re: Claude Opus 5

#183

Wow, 30% on ARC-AGI-3 for $20k total. Huge jump from GPT-5.6's 7.8% at $20k per task. I continue to believe ARC-AGI measures something different and important compared to other benchmarks.

seeing a jump this big is not a great sign for the continuing value of a benchmark

Re: Claude Opus 5

#184
post #164

Their communication is confusing. They say "Opus 5 is not more capable overall than Fable 5", but their blog post proceeds to list how much better Opus 5 is than Fable 5 on __most__ benchmarks listed. Then system card goes on to "Its AI R&D capabilities are comparable to those of Claude Mythos 5", which is supposed to be fable minus restrictions.

Capable in term of AI R&D, not capable in terms of hacking (which caused all the Fable drama.) But agree, confusing wording.

Re: Claude Opus 5

#185
post #89
post #26

Earlier quoted context omitted.

Funny that a company selling an AI software developer can't use it to fix their infra. Fixing those issues still requires humans.

How many companies at the size of Anthropic can serve the amount of traffic and manage the amount of compute they have?

How many compagnies can manage a mostly stateless workload at "whatever-the-scale-because-it-does-not-matter-because-stateless" ? Lots of people can do that. Massive amount of people can do that.

Re: Claude Opus 5

#186

Earlier quoted context omitted.

Go read the safeguards section in the report and you will realize why that is. These models are heavily as safeguarded and that was the initial reason why they said they couldn't and haven't released Mythos because that model is the one without the safeguards. OpenAI is did the same thing when they announced a model without safeguards broken into HuggingFace servers.

Yes, this makes a lot of sense, but it’s just very amusing to see. 2 months ago, the world was about to end, now not so much.

Do you have an example of the "doomsday marketing" you're referring to?

Re: Claude Opus 5

#187

Is Fable 5 just Opus 5 with some additional long-context management modifications for extended self-directed work? Or are they actually truly different models?

based on pricing I think it's safe to say they're different. why would they charge half price when fable has been very popular?

Re: Claude Opus 5

#188
im excited that cad and object=>cad is getting into the test tasks

i guess the next stuff will be tool use for the rest of what cad does in assemblies and simulation?

itd be fun to try to set up a 3d printer as part of a feedback loop, and see what a model can build.

the automated test harness for physical stuff seems a bit beyond reach still

Re: Claude Opus 5

#189
My thoughts: fable is the bigger model. Opus is distilled from it but since it is smaller it doesn’t need the online classifiers. Though benchmarks show Opus to be near Fable level, I think it’s nowhere near Mythos (fable without safeguards).

Re: Claude Opus 5

#190
Seems really good so far using it in Claude Code CLI - it gave me a new flag when I asked a question:

"I don't have a reliable way to read that number, so I'd be guessing if I gave you one — and this is exactly the kind of question where a confident guess is worse than none.

What I can tell you is what I actually observe:"

I really like this update - gave me a clear sense of the facts but didn't give me a guess just for the sake of guessing.

One oddity is that it appears to only have a 200K context window right now via CC. Hopefully the 1M version will appear soon!

Post reply on HN