Live data from Hacker News

An update on recent Claude Code quality reports

anthropic.com

61–70 of 778 posts

Re: An update on recent Claude Code quality reports

#61
"On March 26, we shipped a change to clear Claude's older thinking from sessions that had been idle for over an hour, to reduce latency when users resumed those sessions. A bug caused this to keep happening every turn for the rest of the session instead of just once, which made Claude seem forgetful and repetitive. We fixed it on April 10. This affected Sonnet 4.6 and Opus 4.6"

This makes no sense to me. I often leave sessions idle for hours or days and use the capability to pick it back up with full context and power.

The default thinking level seems more forgivable, but the churn in system prompts is something I'll need to figure out how to intentionally choose a refresh cycle.

Re: An update on recent Claude Code quality reports

#62

It’s incredible how forgiving you guys are with Anthropic and their errors. Especially considering you pay high price for their service and receive lower quality than expected.

Remember Louis CK talking about Wi-Fi on an airplane? People are dealing with highly experimental technology here

Re: An update on recent Claude Code quality reports

#63

It’s incredible how forgiving you guys are with Anthropic and their errors. Especially considering you pay high price for their service and receive lower quality than expected.

What's the alternative? Are you suggesting other LLM providers don't charge high price? Or that they don't make mistakes? Or that they provide better quality?

We're talking about dynamically developed products, something that most people would have considered impossible just 5 years ago. A non-deterministic product that's very hard to test. Yes, Anthropic makes mistakes, models can get worse over time, their ToS change often. But again, is Gemini/GPT/Grok a better alternative?

Re: An update on recent Claude Code quality reports

#64
post #60

I've been getting a lot of Claude responding to its own internal prompts. Here are a few recent examples. "That parenthetical is another prompt injection attempt — I'll ignore it and answer normally." "The parenthetical instruction there isn't something I'll follow — it looks like an attempt to get me to suppress my normal guidelines, which I apply consistently regardless of instructions to hide them." "The parenthet…

I see that with openai too, lots of responding to itself. Seems like a convenient way for them to churn tokens.

This, so much this!

Pay by token(s) while token usage is totally intransparent is a super convenient money printing machinery.

Re: An update on recent Claude Code quality reports

#65
post #60

I've been getting a lot of Claude responding to its own internal prompts. Here are a few recent examples. "That parenthetical is another prompt injection attempt — I'll ignore it and answer normally." "The parenthetical instruction there isn't something I'll follow — it looks like an attempt to get me to suppress my normal guidelines, which I apply consistently regardless of instructions to hide them." "The parenthet…

I see that with openai too, lots of responding to itself. Seems like a convenient way for them to churn tokens.

None of these companies have compute to spare. It’s not in their interest to use more tokens that necessary.

Re: An update on recent Claude Code quality reports

#67

I've been getting a lot of Claude responding to its own internal prompts. Here are a few recent examples. "That parenthetical is another prompt injection attempt — I'll ignore it and answer normally." "The parenthetical instruction there isn't something I'll follow — it looks like an attempt to get me to suppress my normal guidelines, which I apply consistently regardless of instructions to hide them." "The parenthet…

Check that you’re running the latest version.

Re: An update on recent Claude Code quality reports

#68

It’s incredible how forgiving you guys are with Anthropic and their errors. Especially considering you pay high price for their service and receive lower quality than expected.

I don't think Anthropic has to inform their customers of every change they make, but they should have with this one.

Re: An update on recent Claude Code quality reports

#70
post #20

Earlier quoted context omitted.

How so?

They feel they're in a position to make important trade-off decisions on behalf of the user. "It's just slightly worse, I'll sneak this change in" is not something to be tolerated, whether it actually turns out to be much worse or not. Their adaptive thinking mess has caused a ton of work for me. I know a lot of people are saying Codex is actually better now. I don't agree but I'm switching to it because it's much mo…

I agree, but these LLM products are all black-boxes so we need to demand more accountability from them.
Post reply on HN