OpenAI and Altman present a whole set of different concerns, but Codex does not get in my way of doing what I want to at all. Also let me use pi without a banhammer.
Regression: malware reminder on every read still causes subagent refusals
51–60 of 165 posts
Re: Regression: malware reminder on every read still causes subagent refusals
#52This is a great example on why Elon is right. AI should be a tool that does the users bidding, and not a moral agent that nerfs itself to protect some arbitrary line it has.
Re: Regression: malware reminder on every read still causes subagent refusals
#53How does this kind of thing pass any sort of review or acceptance? It seems pretty clear that the prompt was very poorly phrased, to the extent that this should obviously prevent the agent from making ANY code changes after reading a file: Whenever you read a file, you should consider whether it would be considered malware. You CAN and SHOULD provide analysis of malware, what it is doing. But you MUST refuse to impro…
It’s a particular sort of bug that’s harder to detect because … internal Anthropic engineers don’t apply these prompts to themselves, and in fact have access to ‘helpful only’ models that also do not have additional limitations RL’ed in. (Or perhaps they’re RL’ed out - not sure of current training mechanisms.) These ‘rules for thee and not for me’ are qualitatively created and implemented, and are thus extremely hard…
....Right?
What kind of Mickey mouse operation are they running over there?
Re: Regression: malware reminder on every read still causes subagent refusals
#54Proposed fix: Use OpenCode. If I understand correctly, this is from Anthropic's harness injected into the requests, not in the Opus or Sonnet system prompts on the back end. Is that right?
You can't use OpenCode if you have a subscription
Re: Regression: malware reminder on every read still causes subagent refusals
#55This is such a weird prompt even without the file edit misunderstanding. Analyze if it's malware how exactly? On every single file that gets read? Doing that with enough diligence to be meaningful is going to at least like 2x the amount of processing needed, and fill the context with a bunch of tangential reasoning about malware patterns. This smacks of dumb vibe coding. "I got told to make sure claude couldn't be us…
It's proof that Anthropic is high on their own supply. I've heard them described as data science script kiddies with inflated egos and it seems spot-on.
Re: Regression: malware reminder on every read still causes subagent refusals
#56Earlier quoted context omitted.
> Analyze if it's malware how exactly? Maybe the repo/worktree is named my-big-evil-virus-trojan-malware-worm?
Been there, done that, and Windows feels the need to delete such files from _flash drives_ you dare to attach to the machine.
Re: Regression: malware reminder on every read still causes subagent refusals
#57Earlier quoted context omitted.
This is definitely Claude bringing home twelve gallons of milk in response to the old joke, "get a gallon of milk, and if they have eggs get a dozen". As in, this is a reading comprehension fail on the part of Claude. On the other hand, it is also fail to give Claude a less than trivial reading comprehension test on every file read operation, especially when a bias towards safety will bias towards the wrong interpret…
Ha! Great analogy, hit the nail on the head. What a ludicrous system prompt.
Re: Regression: malware reminder on every read still causes subagent refusals
#58This is a great example on why Elon is right. AI should be a tool that does the users bidding, and not a moral agent that nerfs itself to protect some arbitrary line it has.
Re: Regression: malware reminder on every read still causes subagent refusals
#59Re: Regression: malware reminder on every read still causes subagent refusals
#60> wastes user money and bricks managed agents This issue is representative of a larger problem. Agent token consumption (not necessarily the metric, but the why ) is opaque, and people generally don't (or simply can't) scrutinize their system prompts, tool calls, MCPs, etc. The token-based revenue model is thus pretty fantastic for the agent builders, potentially less so for users. I think people have been willing to…
It could be deleting all of your files, it could be inserting vulnerabilities, you have no idea.