Live data from Hacker News

Elevated errors on Claude Opus 5

status.claude.com

1–10 of 79 posts

Re: Elevated errors on Claude Opus 5

#4

I would say "Elected errors _in_ Claude Opus 5" wouldn't be incorrect either.. Opus 5 isn't very reliable for coding and introduces a lot of regressions every single time I use it. Do you have the same experiences?

Do you not have unit tests, or how does it introduce regressions? You can tell it how to run the test suite in CLAUDE.md

Re: Elevated errors on Claude Opus 5

#5
post #4

I would say "Elected errors _in_ Claude Opus 5" wouldn't be incorrect either.. Opus 5 isn't very reliable for coding and introduces a lot of regressions every single time I use it. Do you have the same experiences?

Do you not have unit tests, or how does it introduce regressions? You can tell it how to run the test suite in CLAUDE.md

I detect the regression already in planning with Opus 5, so I do not let Opus 5 implement anything. But it is a waste of time and tokens! Does planning with Opus 5 works out for you?

Re: Elevated errors on Claude Opus 5

#6
post #4

I would say "Elected errors _in_ Claude Opus 5" wouldn't be incorrect either.. Opus 5 isn't very reliable for coding and introduces a lot of regressions every single time I use it. Do you have the same experiences?

Do you not have unit tests, or how does it introduce regressions? You can tell it how to run the test suite in CLAUDE.md

Opus 5 tries to modify the unit tests as a cover to its own regressions - thinking its own logic is correct and the test must be wrongly specified

Re: Elevated errors on Claude Opus 5

#7
post #4

I would say "Elected errors _in_ Claude Opus 5" wouldn't be incorrect either.. Opus 5 isn't very reliable for coding and introduces a lot of regressions every single time I use it. Do you have the same experiences?

Do you not have unit tests, or how does it introduce regressions? You can tell it how to run the test suite in CLAUDE.md

It is a well known fact that projects with unit tests never have regressions.

Re: Elevated errors on Claude Opus 5

#8

I would say "Elected errors _in_ Claude Opus 5" wouldn't be incorrect either.. Opus 5 isn't very reliable for coding and introduces a lot of regressions every single time I use it. Do you have the same experiences?

Opus 5 isn't very reliable for coding and introduces a lot of regressions every single time I use it.

And GPT 5.6 Sol over engineers just about everything. No LLM is perfect, its about learning the issues with each LLM and figuring out if you can live with it. Knowledge means that you can anticipate if it tries to pull something funny, and harness it against that behavior.

Re: Elevated errors on Claude Opus 5

#9

I would say "Elected errors _in_ Claude Opus 5" wouldn't be incorrect either.. Opus 5 isn't very reliable for coding and introduces a lot of regressions every single time I use it. Do you have the same experiences?

Maybe it's my harness but I haven't seen it introducing regressions.

Re: Elevated errors on Claude Opus 5

#10

I would say "Elected errors _in_ Claude Opus 5" wouldn't be incorrect either.. Opus 5 isn't very reliable for coding and introduces a lot of regressions every single time I use it. Do you have the same experiences?

Opus 5 isn't very reliable for coding and introduces a lot of regressions every single time I use it. And GPT 5.6 Sol over engineers just about everything. No LLM is perfect, its about learning the issues with each LLM and figuring out if you can live with it. Knowledge means that you can anticipate if it tries to pull something funny, and harness it against that behavior.

this would be _Great_ advice if you owned your own LLM and your knowledge was trapped in Amber because you were satisfied.

It's horrible advice given what we've seen consistent: changing alignments, changing guardrails, changing system prompts, changing inference priorities, etc.

Anyone who relies on these for their work product is chaining themselves to a matrix multiple of indetermintism.

Post reply on HN