Live data from Hacker News

Issue: Claude Code is unusable for complex engineering tasks with Feb updates

github.com

741–750 of 829 posts

Re: Issue: Claude Code is unusable for complex engineering tasks with Feb updates

#741
post #590

Earlier quoted context omitted.

If distilled models were commercially banned they'd probably be willing to show the thinking again.

How do you think such a ban should work? Do you not see that the next (or previous) logical step would be a "commercial ban" of frontier models, all "distilled" from an enormous amount of copyrighted material?

I'm not arguing the merits of such a ban, I'm simply stating a fact - that thinking transcripts likely won't return until such a ban is in place.

Re: Issue: Claude Code is unusable for complex engineering tasks with Feb updates

#743
post #445

I put together a quick audit to check for "early landing" messages[1] using jq, ripgrep, and the messages[2] flagged in the stop guard script. I have noticed a trend in these sessions asking more and more about calling it a day, "it's getting late," and other phrases. I sort of assumed it was some kind of "load shedding" on Anthropic's side. My audit of 80 sessions was interesting. Sorry, I won't share details, but I…

As a negative example, my audit of 31 sessions was uninteresting. I had one matching entry, where I had pasted a long list of console errors into Claude and it identified a few as pre-existing and asked me to get more information for follow-up analysis.

I wonder if it comes down to prompting—maybe by introducing these "golden rules" OP mentions in their CLAUDE.md, they're actually "priming" Claude to think about these stop phrases and introduce them proactively.

Do you have a CLAUDE.md file? What does it contain?

Re: Issue: Claude Code is unusable for complex engineering tasks with Feb updates

#744
post #179

I'm the author of the report in there. The stop-phrase-guard didn't get attached but here it is: https://gist.github.com/benvanik/ee00bd1b6c9154d6545c63e06a3... You can watch for these yourself - they are strong indicators of shallow thinking. If you still have logs from Jan/Feb you can point claude at that issue and have it go look for the same things (read:edit ratio shifts, thinking character shifts before the red…

The "this test failure is preexisting so I'm going to ignore it" thing has been happening a lot for me lately, it's so annoying. Unless it makes a change and then immediately runs tests and it's obvious from the name/contents that the failing test is directly related to the change that was made it will ignore it and not try to fix.

I will note that this "out" that Claude takes was a) less frequent in Opus 4.5 and that time frame and b) notably not something that Codex does.

I don't trust the code that Claude writes at all, if I have to use it (they gave me a free month recently, so I use it...) I not only review it carefully but have Codex do a thorough review.

Claude "cheats" and leaves hacks and has Dunning-Kruger.

All of this is very exhausting. I am enjoying writing my own code with these tools (to get long running personal projects out the door) but the effect that these tools are having on teams is terrifyingly corrosive and it's making me want to take an early retirement from the profession.

Yes we can write a lot of code quickly. But at what cost? And what even use is all this code now anyways?

Re: Issue: Claude Code is unusable for complex engineering tasks with Feb updates

#745
post #616

I appreciate the work done here. Been having this feeling that things have got worse recently but didn't think it could be model related. The most frustrating aspect recently (I have learned and accepted that Claude produces bad code and probably always did, mea culpa) is the non-compliance. Claude is racing away doing its own thing, fixing things i didn't ask, saying the things it broke are nothing to do with it, et…

I am still on an old version of CC on one machine, but the results are the same. More difficulty keeping it on track, convincing it timelines I suggest are correct etc. For example I had a deploy fail, and it would not believe that the new logs were not from a previous deploy. It was adamant it had fixed the issue, so the logs must be old logs.

I was using web UI last night and it was unable to understand basic aspects of the task. Haven't seen it perform this badly since I began using two years ago.

Was trying to track token usage/index with Cursor, and was unable to understand that running `find` wouldn't show what was in Cursor index. Multiple times.

Re: Issue: Claude Code is unusable for complex engineering tasks with Feb updates

#746

Hey all, Boris from the Claude Code team here. I just responded on the issue, and cross-posting here for input. --- Hi, thanks for the detailed analysis. Before I keep going, I wanted to say I appreciate the depth of thinking & care that went into this. There's a lot here, I will try to break it down a bit. These are the two core things happening: > `redact-thinking-2026-02-12` This beta header hides thinking from th…

Didn’t ULTRATHINK get deprecated? Last time I typed it I got a warning.

Re: Issue: Claude Code is unusable for complex engineering tasks with Feb updates

#747
post #707

Hey all, Boris from the Claude Code team here. I just responded on the issue, and cross-posting here for input. --- Hi, thanks for the detailed analysis. Before I keep going, I wanted to say I appreciate the depth of thinking & care that went into this. There's a lot here, I will try to break it down a bit. These are the two core things happening: > `redact-thinking-2026-02-12` This beta header hides thinking from th…

This is confusing. ULTRATHINK is a step below /effort max? ULTRATHINK triggers high effort. /effort max is above high. Calling it ULTRATHINK sounds like it would be the highest mode. If someone has max set and types ULTRATHINK, they're lowering their effort for that turn. For anyone reading this trying to fix the quality issues, here's what I landed on in ~/.claude/settings.json: { "env": { "CLAUDE_CODE_EFFORT_LEVEL"…

Thanks for sharing. Have you experienced noticeable impact to your usage rate?

Re: Issue: Claude Code is unusable for complex engineering tasks with Feb updates

#748
post #650

Earlier quoted context omitted.

i mean you could just search up "is Anthropic making profit" and most sources will say no. There's this one source on Reddit which calculated that Anthropic has been subsidizing their costs by 32x

I really wonder about this. Is it so bad that they cannot even disclose it? not even an optimistic lie in the ballpark of reality? it's not like they haven't been found cooking the truth repeatedly. I look at the output of Kimi and the costs of running inference on it that i can replicate, and it isn't that bad, although admittedly i don't have to worry anywhere near as much about scaling it and about having to dedic…

the biggest red flag I see is this: https://youtu.be/iOyFja87uyw?si=5INnIG1kZI0AbCGa

tldr: they are trying hard to change S&P500 inclusion rules so that they dont have to wait 12months after going public so they can list mega-ipo asap in force index funds to buy a portion (presumably before revenue exponential growth settles and profits start tanking due to opensource catching up). They know something that we dont.

btw if they are public and part of S&P500 then potentially they'll be a candidate for a bailout.

Re: Issue: Claude Code is unusable for complex engineering tasks with Feb updates

#749

Hey all, Boris from the Claude Code team here. I just responded on the issue, and cross-posting here for input. --- Hi, thanks for the detailed analysis. Before I keep going, I wanted to say I appreciate the depth of thinking & care that went into this. There's a lot here, I will try to break it down a bit. These are the two core things happening: > `redact-thinking-2026-02-12` This beta header hides thinking from th…

There's been more going on than just the default to medium level thinking - I'll echo what others are saying, even on high effort there's been a very significant increase in "rush to completion" behavior.

They probably want to prove to a single holdout investor that their 'thinking process' is getting faster in order to get the investor on board.

Re: Issue: Claude Code is unusable for complex engineering tasks with Feb updates

#750
post #597

Earlier quoted context omitted.

Anthropic's position is that thinking tokens aren't actually faithful to the internal logic that the LLM is using, which may be one reason why they started to exclude them: https://www.anthropic.com/research/reasoning-models-dont-say...

That probably matters for some scenarios, but I have yet to find one where thinking tokens didn't hint at the root cause of the failure. All of my unsupervised worker agents have sidecars that inject messages when thinking tokens match some heuristics. For example, any time opus says "pragmatic", its instant Esc Esc > "Pragmatic fix is always wrong, do the Correct fix", also whenever "pre-existing issue" appears (it'…

I had some interesting experience to the opposite last night, one of my tests has been failing for a long time, something to do with dbus interacting with Qt segfaulting pytest. Been ignoring it for a long time, finally asked claude code to just remove the problematic test. Come back a few minutes later to find claude burning tokens repeatedly trying and failing to fix it. "Actually on second thought, it would be better to fix this test."

Match my vibes, claude. The application doesn't crash, so just delete that test!

Post reply on HN