Live data from Hacker News

Elevated errors on Claude Opus 5

status.claude.com

51–60 of 79 posts

Re: Elevated errors on Claude Opus 5

#51
post #24

Lots of errors. Opus 5 is also giving me many more hallucinations, including things that aren't even in the right territory. It's also telling me that it's making many mistakes, and the language feels off-kilter as if it's not using typical clear phrases.

Yeah I noticed informing me of mistakes it made during sessions. It felt really off when it informed me of a mistake it almost made but caught it before it landed.

Re: Elevated errors on Claude Opus 5

#54
post #27

Earlier quoted context omitted.

I remember when 4.7 and 4.8 were released and people were asking what's wrong with them and 4.6 is the best. But yes, I also think it's not the greatest model for programming. On the other hand, for agentic tasks that are not programming related it's hard to beat Opus 4.8. It can try different things and pivot even when the user is not great with prompting. 5.0 seems to not be worse, but definitely wastes more tokens…

4.6 was better in some way that I can’t put my finger on. None of the models since have been able to reproduce its quality of output for me.

is, not was. Thankfully 4.6 is still being served by Anthropic.

Re: Elevated errors on Claude Opus 5

#59
post #24

Lots of errors. Opus 5 is also giving me many more hallucinations, including things that aren't even in the right territory. It's also telling me that it's making many mistakes, and the language feels off-kilter as if it's not using typical clear phrases.

Yes, it’s way off field in many things, gets into weird minutia without seeing a way out, and it’s often seeing a clear sequence of work but then halts on a statement like “ok I’m going to start now.” Then after expiring the cache when I notice and ask why didn’t it the response is “no reason starting now!”

I see this behavior constantly in 5 - the quality of opus and fable have degraded constantly since 4.6 was such a riotous success

Re: Elevated errors on Claude Opus 5

#60

I would say "Elected errors _in_ Claude Opus 5" wouldn't be incorrect either.. Opus 5 isn't very reliable for coding and introduces a lot of regressions every single time I use it. Do you have the same experiences?

Yes I had a similar experience with Opus 5. It is very token efficient, fast, and gets reasonable part of the work right but makes a LOT of mistakes. In a month+ use of Fable completed each task without ANY errors. Opus could not complete a single of ~5 tasks without some issue or the other - either not getting it fully right or actually introducing regressions. To their credit it was able to catch regressions and fix competently. It seems like a pre Opus 4.6 model in terms of reliability with a lot more power and spiky intelligence. When it gets things right it's powerful and efficient but without reliability I had to 'downgrade' to Opus 4.8 forcibly (since it was not a default option on claude code). I really miss Fable on the pro plan and will likely churn to K3.
Post reply on HN