Live data from Hacker News

Claude Opus 5

anthropic.com

621–630 of 1001 posts

Re: Claude Opus 5

#621
post #38

> Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source-code vulnerability discovery at all access levels. This means that the model can support defensive cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively. Okay so it’s worse than Opus 4.8 for my purposes I guess?

Presumably it drops back to 4.8 in those cases so it's not really worse

If it switches mid conversation, this is a massive increase in token consumption because it has to re-read your conversation into cache, right?

Re: Claude Opus 5

#622
post #617

Earlier quoted context omitted.

It's not that important in most cases, but yes, on the aggregate, it's a concern

It's a binary rule at our firm. We trust bedrock but not mantle for similar reasons

Yes, that's valid, I'm sure it's a dealbraker in some situations.

Re: Claude Opus 5

#623
I had a moderately complex review in a large C/C++ codebase that Codex/GPT-5.6-sol already cleaned up so I threw it at Opus 5. 4 errors found. That seemed odd, so I handed it back to GPT. All were false. Opus doesn't seem to look at the wider context and understand which functions were called in certain contexts. I gave GPT's analysis back to Opus and it admitted its mistake. Maybe it's good for writing code, but as far as analysis it seems like it needs some work.

Re: Claude Opus 5

#624

I had a moderately complex review in a large C/C++ codebase that Codex/GPT-5.6-sol already cleaned up so I threw it at Opus 5. 4 errors found. That seemed odd, so I handed it back to GPT. All were false. Opus doesn't seem to look at the wider context and understand which functions were called in certain contexts. I gave GPT's analysis back to Opus and it admitted its mistake. Maybe it's good for writing code, but as…

That has always been the major strength of GPT, that's the model you use for checking. It often nearly isn't as good for creation though.

Re: Claude Opus 5

#625

I've yet to understand why they call a 190 page PDF a "card". Calling something a card invokes a small, quick rundown of pertinent details, not every single possible detail.

Because "model card" is a set phrase, it's a concept. It originates from a time when they were shorter. Like datasheets, even if it's not literally a sheet. They could say "tech report" but model card makes it clear that it's a specific kind of tech report.

[dead]

Re: Claude Opus 5

#626
post #529

Earlier quoted context omitted.

Fable might be using those phrases less, but its writing is still terrible and exhausting to read.

Agreed. I’ve interestingly found 5.6 sol to produce much better writing, and it can generally cut to the point much more effectively.

Neither Claude nor GPT are acceptable for writing English text. Personally I have found Gemini to be far better, and that is really all I use it for.

Re: Claude Opus 5

#627
post #517

Wait, 30% on ARC-AGI-3! I definitely didn't expect that jump so soon. Are there any rumors of what they are changing in architecture that is leading to this?

30% on ARC-AGI-3 is the first two puzzles. It cost $20000 in tokens to do that. That is a terrible result that doesn't imply anything.

This happened before for arc agi 1 and 2 great gains but high costs then slowly but surely the price dropped for less than a dollar per task and got saturated.

Re: Claude Opus 5

#630
post #617

Earlier quoted context omitted.

It's not that important in most cases, but yes, on the aggregate, it's a concern

It's a binary rule at our firm. We trust bedrock but not mantle for similar reasons

why not mantle? - I was confused why mantle exists over normal bedrock
Post reply on HN