Live data from Hacker News

An update on recent Claude Code quality reports

anthropic.com

1–10 of 778 posts

Re: An update on recent Claude Code quality reports

#2
1. They changed the default in March from high to medium, however Claude Code still showed high (took 1 month 3 days to notice and remediate)

2. Old sessions had the thinking tokens stripped, resuming the session made Claude stupid (took 15 days to notice and remediate)

3. System prompt to make Claude less verbose reducing coding quality (4 days - better)

All this to say... the experience of suspecting a model is getting worse while Anthropic publicly gaslights their user-base: "we never degrade model performance" is frustrating.

Yes, models are complex and deploying them at scale given their usage uptick is hard. It's clear they are playing with too many independent variables simultaneously.

However you are obligated to communicate honestly to your users to match expectations. Am I being A/B tested? When was the date of the last system prompt change? I don't need to know what changed, just that it did, etc.

Doing this proactively would certainly match expectations for a fast-moving product like this.

Re: An update on recent Claude Code quality reports

#3
post #2

1. They changed the default in March from high to medium, however Claude Code still showed high (took 1 month 3 days to notice and remediate) 2. Old sessions had the thinking tokens stripped, resuming the session made Claude stupid (took 15 days to notice and remediate) 3. System prompt to make Claude less verbose reducing coding quality (4 days - better) All this to say... the experience of suspecting a model is get…

[deleted]

Re: An update on recent Claude Code quality reports

#4
The issue making Claude just not do any work was infuriating to say the least. I already ran at medium thinking level so was never impacted, but having to constantly go "okay now do X like you said" was annoying.

Again goes back to the "intern" analogy people like to make.

Re: An update on recent Claude Code quality reports

#5
Wow, bad enough for them to actually publish something and not cryptic tweets from employees.

Damage is done for me though. Even just one of these things (messing with adaptive thinking) is enough for me to not trust them anymore. And then their A/B testing this week on pricing.

Re: An update on recent Claude Code quality reports

#6
post #2

1. They changed the default in March from high to medium, however Claude Code still showed high (took 1 month 3 days to notice and remediate) 2. Old sessions had the thinking tokens stripped, resuming the session made Claude stupid (took 15 days to notice and remediate) 3. System prompt to make Claude less verbose reducing coding quality (4 days - better) All this to say... the experience of suspecting a model is get…

None of these problems equate to degrading model performance. Completely different team. Degraded CC harness, sure.

Re: An update on recent Claude Code quality reports

#8
> On April 16, we added a system prompt instruction to reduce verbosity. In combination with other prompt changes, it hurt coding quality, and was reverted on April 20. This impacted Sonnet 4.6, Opus 4.6, and Opus 4.7.

Claude caveman in the system prompt confirmed?

Re: An update on recent Claude Code quality reports

#9
post #2

1. They changed the default in March from high to medium, however Claude Code still showed high (took 1 month 3 days to notice and remediate) 2. Old sessions had the thinking tokens stripped, resuming the session made Claude stupid (took 15 days to notice and remediate) 3. System prompt to make Claude less verbose reducing coding quality (4 days - better) All this to say... the experience of suspecting a model is get…

> Anthropic publicly gaslights their user-base: "we never degrade model performance" is frustrating.

They're not gaslighting anyone here: they're very clear that the model itself, as in Opus 4.7, was not degraded in any way (i.e. if you take them at their word, they do not drop to lower quantisations of Claude during peak load).

However, the infrastructure around it - Claude Code, etc - is very much subject to change, and I agree that they should manage these changes better and ensure that they are well-communicated.

Post reply on HN