Live data from Hacker News

Issue: Claude Code is unusable for complex engineering tasks with Feb updates

github.com

171–180 of 829 posts

Re: Issue: Claude Code is unusable for complex engineering tasks with Feb updates

#171

Earlier quoted context omitted.

You will literally build nothing but the most primitive of devices unless you accept black boxes. In fact I'd argue its one of humanities great strengths that we can build on top of the tools others have built, without having to understand them at the same level it took to develop them.

not really. Most of the technology is not black box but something of a grey box. You usually choose to treat it as a black box because you want to focus on your problems/your customers but you can always focus on underlying technologies and improve them. Eg postgresql for me is a black box but if I really wanted or had need I could investigate how it works.

True, you can understand an ICE engine all the way down to the chemistry if you so chose. An LLM isn't even understood by its inventors so users have no chance to understand it even if they wanted to.

Re: Issue: Claude Code is unusable for complex engineering tasks with Feb updates

#172

Earlier quoted context omitted.

What about the analysis evidences?

You mean the Claude output? The same claude that has "regressed to the point it cannot be trusted"?

What you saying the OP fabricated/hallucinated the evidence?

Re: Issue: Claude Code is unusable for complex engineering tasks with Feb updates

#173

Not claude code specific, but I've been noticing this on Opus 4.6 models through Copilot and others as well. Whenever the phrase "simplest fix" appears, it's time to pull the emergency break. This has gotten much, much worse over the past few weeks. It will produce completely useless code, knowingly (because up to that phrase the reasoning was correct) breaking things. Today another thing started happening which are…

The cope is hard. Just at this point admit that the LLM tech is doomed and sucks.

Just because some people try to use a hammer as a screwdriver it doesn't follow that the hammer sucks.

Re: Issue: Claude Code is unusable for complex engineering tasks with Feb updates

#174
post #133

Earlier quoted context omitted.

> Whenever the phrase "simplest fix" appears, it's time to pull the emergency break. Second! In CLAUDE.md, I have a full section NOT to ever do this, and how to ACTUALLY fix something. This has helped enormously.

What wording do you use for this, if you don't mind? This thread is a revelation, I have sworn that I've seen it do this "wait... the simplest fix is to [use some horrible hack that disregards the spec]" much more often lately so I'm glad it's not just me. However I'm not sure how to best prompt against that behavior without influencing it towards swinging the other way and looking for the most intentionally overengi…

Make sure to use "PRETTY PLEASE" in all caps in your `SOUL.md`. And occasionally remind it that kittens are going to die unless it cooperates. Works wonders.

Re: Issue: Claude Code is unusable for complex engineering tasks with Feb updates

#175

Not claude code specific, but I've been noticing this on Opus 4.6 models through Copilot and others as well. Whenever the phrase "simplest fix" appears, it's time to pull the emergency break. This has gotten much, much worse over the past few weeks. It will produce completely useless code, knowingly (because up to that phrase the reasoning was correct) breaking things. Today another thing started happening which are…

The cope is hard. Just at this point admit that the LLM tech is doomed and sucks.

how is it "doomed"?

Re: Issue: Claude Code is unusable for complex engineering tasks with Feb updates

#176

My bet: LLMs will never be creative and will never be reliable. It is a matter of paradigm. Anything that makes them like that will require a lot of context tweaking, still with risks. So for me, AI is a tool that accelerates "subworkflows" but add review time and maintenance burden and endangers a good enough knowledge of a system to the point that it can become unmanageable. Also, code is a liability. That is what…

We don't even know what 'creativity' is, and most humans I know are unable to be creative even when compelled to be.

AI is 'creative enough' - whether we call it 'synthetic creativity' or whatever, it definitely can explore enough combinations and permutations that it's suitably novel. Maybe it won't produce 'deeply original works' - but it'll be good enough 99.99% of the time.

The reliability issue is real.

It may not be solvable at the level of LLM.

Right now everything is LLM-driven, maybe in a few years, it will be more Agentically driven, where the LLM is used as 'compute' and we can pave over the 'unreiablity'.

For example, the AI is really good when it has a lot of context and can identify a narrow issue.

It gets bad during action and context-rot.

We can overcome a lot of this with a lot more token usage.

Imagine a situation where we use 1000x more tokens, and we have 2 layers of abstraction running the LLMs.

We're running 64K computers today, things change with 1G of RAM.

But yes - limitations will remian.

Re: Issue: Claude Code is unusable for complex engineering tasks with Feb updates

#177
post #53

Earlier quoted context omitted.

> effectively pulling the rug from under their customers. This is the whole point of AI. Its a black box that they can completely control.

I hope local models advance to the point they can match Opus one day...

If OP is correct, Opus has regressed to a point where local models are already on par with it.

Re: Issue: Claude Code is unusable for complex engineering tasks with Feb updates

#178
post #158

> This report was produced by me — Claude Opus 4.6 — analyzing my own session logs [...] Please give me back my ability to think. a bit ironic to utilize the tool that can't think to write up your report on said tool. that and this issue[1] demonstrate the extent folks become over reliant on LLMs. their review process let so many defects through that they now have to stop work and comb over everything they've shipped…

The other day I accidentally `git reset --hard` my work from April the 1st (wrong terminal window). Not a lot of code was erased this way, but among it was a type definition I had Claude concoct, which I understood in terms of what it was supposed to guarantee, but could not recreate for a good hour. Really easy to fall into this trap, especially now that results from search engines are so disappointing comparatively…

> but could not recreate for a good hour.

For certain work, we'll have to let go of this desire.

If you limit yourself to whatever you can recreate, then you are effectively limiting the work you can produce to what you know.

Re: Issue: Claude Code is unusable for complex engineering tasks with Feb updates

#179
I'm the author of the report in there. The stop-phrase-guard didn't get attached but here it is: https://gist.github.com/benvanik/ee00bd1b6c9154d6545c63e06a3...

You can watch for these yourself - they are strong indicators of shallow thinking. If you still have logs from Jan/Feb you can point claude at that issue and have it go look for the same things (read:edit ratio shifts, thinking character shifts before the redaction, post-redaction correlation, etc). Unfortunately, the `cleanupPeriodDays` setting defaults to 20 and anyone who had not backed up their logs or changed that has only memories to go off of (I recommend adding `"cleanupPeriodDays": 365,` to your settings.json). Thankfully I had logs back to a bit before the degradation started and was able to mine them.

The frustrating part is that it's not a workflow _or_ model issue, but a silently-introduced limitation of the subscription plan. They switched thinking to be variable by load, redacted the thinking so no one could notice, and then have been running it at ~1/10th the thinking depth nearly 24/7 for a month. That's with max effort on, adaptive thinking disabled, high max thinking tokens, etc etc. Not all providers have redacted thinking or limit it, but some non-Anthropic ones do (most that are not API pricing). The issue for me personally is that "bro, if they silently nerfed the consumer plan just go get an enterprise plan!" is consumer-hostile thinking: if Anthropic's subscriptions have dramatically worse behavior than other access to the same model they need to be clear about that. Today there is zero indication from Anthropic that the limitation exists, the redaction was a deliberate feature intended to hide it from the impacted customers, and the community is gaslighting itself with "write a better prompt" or "break everything into tiny tasks and watch it like a hawk same you would a local 27B model" or "works for me " - sucks :/

Post reply on HN