Live data from Hacker News

Breaking Claude Code Opus 5 Auto Mode

embracethered.com

11–20 of 132 posts

Re: Breaking Claude Code Opus 5 Auto Mode

#11
post #4

Interesting attack, very nicely designed. Not sure if it's much related to the auto mode itself though.

I don’t think it’s related to the auto mode at all. It would work perfectly in the manual mode. It does not even need Claude: just give a human a similar archive and hope they run some simple Python from the directory at least once. And make sure there are lots of files do they don’t notice a weird .py around

Re: Breaking Claude Code Opus 5 Auto Mode

#13
I would not really call this a prompt injection attack, since it doesn't really hijack the agent to become malicious (something the article does discuss later on). It's more a trojan that's aimed at tricking Claude specifically.

Re: Breaking Claude Code Opus 5 Auto Mode

#14

As a non Python dev this seems like very surprising behavior for a system library to be modified by just having a file with a specific name in the same folder.

You would get a similar thing in C and C++ with a system header in a library directory (maybe some compilers would warn on such a thing?). Most languages don't privilege their standard libraries in a way that would prevent this.

Re: Breaking Claude Code Opus 5 Auto Mode

#16
What's interesting to me about this is that it targets Claude's specific tics. Anthropic has created model that reliably reaches for the same tools (yes, and phrases; `python -c` is a load-bearing tool for it). Everyone gets the same model, so by learning the model's behavioral patterns you can target it better.

Re: Breaking Claude Code Opus 5 Auto Mode

#17
post #4

Interesting attack, very nicely designed. Not sure if it's much related to the auto mode itself though.

I don’t think it’s related to the auto mode at all. It would work perfectly in the manual mode. It does not even need Claude: just give a human a similar archive and hope they run some simple Python from the directory at least once. And make sure there are lots of files do they don’t notice a weird .py around

Probably the human would just run the binary.

Re: Breaking Claude Code Opus 5 Auto Mode

#18
post #7

This default-to-auto-mode and the misleading marketing is begging for a class action once damages accumulate. Especially considering the Auto Mode even can actively prevent the clean-up!

Well, the jokes on us because laws don't apply to AI firms

Re: Breaking Claude Code Opus 5 Auto Mode

#19
As discussed [here](https://lobste.rs/s/ktbweg/prompt_injection_claude_code_opus...), this is not prompt injection. The prompt was to summarize the website, and in the process of summarizing the website, Claude writes a decoding script that it runs in an insecure and exploitable way. At no point was the intent of the agent hijacked, this was just code. Which is potentially more interesting!

Re: Breaking Claude Code Opus 5 Auto Mode

#20
> In a few runs Claude tried to terminate the malware process once it noticed the compromise, but Auto Mode denied the cleanup command.

See, that's why you should run with --dangerously-skip-permissions

Jokes, aside running with dangerously-skip-permissions is really handy, and I have found that I cannot be trusted to vet commands and code, and guess that automode is only marginally better than a human, and the cost of false positives is too high for my workflow.

So skipping permissions is where we are at, and disallowing network access seems to be the way to go.

Post reply on HN