Live data from Hacker News

Issue: Claude Code is unusable for complex engineering tasks with Feb updates

github.com

261–270 of 829 posts

Re: Issue: Claude Code is unusable for complex engineering tasks with Feb updates

#261

Hey all, Boris from the Claude Code team here. I just responded on the issue, and cross-posting here for input. --- Hi, thanks for the detailed analysis. Before I keep going, I wanted to say I appreciate the depth of thinking & care that went into this. There's a lot here, I will try to break it down a bit. These are the two core things happening: > `redact-thinking-2026-02-12` This beta header hides thinking from th…

[flagged]

Technically speaking, models inherently do this - CoT is just output tokens that aren't included in the final response because they're enclosed in tags, and it's the model that decides when to close the tag. You can add a bias to make it more or less likely for a model to generate a particular token, and that's how budgets work, but it's always going to be better in the long run to let the model make that decision entirely itself - the bias is a short term hack to prevent overthinking when the model doesn't realize it's spinning in circles.

Re: Issue: Claude Code is unusable for complex engineering tasks with Feb updates

#263
post #256

Hey all, Boris from the Claude Code team here. I just responded on the issue, and cross-posting here for input. --- Hi, thanks for the detailed analysis. Before I keep going, I wanted to say I appreciate the depth of thinking & care that went into this. There's a lot here, I will try to break it down a bit. These are the two core things happening: > `redact-thinking-2026-02-12` This beta header hides thinking from th…

[flagged]

[dead]

Re: Issue: Claude Code is unusable for complex engineering tasks with Feb updates

#264

Hey all, Boris from the Claude Code team here. I just responded on the issue, and cross-posting here for input. --- Hi, thanks for the detailed analysis. Before I keep going, I wanted to say I appreciate the depth of thinking & care that went into this. There's a lot here, I will try to break it down a bit. These are the two core things happening: > `redact-thinking-2026-02-12` This beta header hides thinking from th…

Thinking time is not the issue. The issue is that Claude does not actually complete tasks. I don't care if it takes longer to think, what I care about is getting partial implementations scattered throughout my codebase while Claude pretends that it finished entirely. You REALLY need to fix this, it's atrocious.

Re: Issue: Claude Code is unusable for complex engineering tasks with Feb updates

#265

Earlier quoted context omitted.

> This beta header hides thinking from the UI, since most people don't look at it. I look at it, and I am very upset that I no longer see it.

There is a setting if you'd like to continue to see it: showThinkingSummaries. See the docs: https://code.claude.com/docs/en/settings#available-settings

> As I noted in the comment,

Piece of free PR advice: this is fine in a nerd fight, but don't do this in comments that represent a company. Just repeat the relevant information.

Re: Issue: Claude Code is unusable for complex engineering tasks with Feb updates

#266

Hey all, Boris from the Claude Code team here. I just responded on the issue, and cross-posting here for input. --- Hi, thanks for the detailed analysis. Before I keep going, I wanted to say I appreciate the depth of thinking & care that went into this. There's a lot here, I will try to break it down a bit. These are the two core things happening: > `redact-thinking-2026-02-12` This beta header hides thinking from th…

I was not aware the default effort had changed to medium until the quality of output nosedived. This cost me perhaps a day of work to rectify. I now ensure effort is set to max and have not had a terrible session since. Please may I have a "always try as hard as you can" mode ?

Re: Issue: Claude Code is unusable for complex engineering tasks with Feb updates

#267

Earlier quoted context omitted.

How should one conduct such a rigourously reproducible experiment when LLMs by nature aren't deterministic and when you don't have access to the model you are comparing to from months ago?

Something like this: https://marginlab.ai/trackers/claude-code/ (see methodology section)

[deleted]

Re: Issue: Claude Code is unusable for complex engineering tasks with Feb updates

#268

Earlier quoted context omitted.

ChatGPT has been doing the same consistently for years. Model starts out smooth, takes a while, and produces good (relatively) results. Within a few weeks, responses start happening much more quickly, at a poorer quality.

people have been complaining about this since GPT-4 and have never been able to provide any evidence (even though they have all their old conversations in their chat history). I think it’s simply new model shininess turning into raised expectations after some amount of time.

I agree with you. I too complain about this same phenomenon with my colleagues, and we always arrive at the same conclusion: it’s probably us just expecting more and more over time.

Re: Issue: Claude Code is unusable for complex engineering tasks with Feb updates

#269
post #256

Hey all, Boris from the Claude Code team here. I just responded on the issue, and cross-posting here for input. --- Hi, thanks for the detailed analysis. Before I keep going, I wanted to say I appreciate the depth of thinking & care that went into this. There's a lot here, I will try to break it down a bit. These are the two core things happening: > `redact-thinking-2026-02-12` This beta header hides thinking from th…

[flagged]

It also completely ignores the increase in behavioral tracking metrics. 68% increase in swearing at the LLM for doing something wrong needs to be addressed and isn't just "you're holding it wrong"

Re: Issue: Claude Code is unusable for complex engineering tasks with Feb updates

#270
post #256

Earlier quoted context omitted.

[flagged]

I’m not sure being confrontational like this really helps your case. There are real people responding, and even if you’re frustrated it doesn’t pay off to take that frustration out on the people willing to help.

Is somebody saying "you're holding it wrong" a "people willing to help"?
Post reply on HN