Live data from Hacker News

Meta Security Researcher's AI Agent Accidentally Deleted Her Emails

au.pcmag.com

21–30 of 68 posts

Re: Meta Security Researcher's AI Agent Accidentally Deleted Her Emails

#23
post #16

Earlier quoted context omitted.

> Yesterday Claude Code made 5 large edits to my codebase in planning mode (claims it used a bash script instead of standard read/write tools so the guardrails didn't trigger) it's why I put agents in containers. More and more my prompts have to tell Claude what I don't want it to do . It's crazy to me I'm arguing with it, having to ask and convince it to do the right things. Regardless of sentience (I'm not touching…

The other day, it was a little like Claude was trying to find a loophole for ignoring my instructions, and being punchy about it: CLAUDE: [...] Did I use npm: yes — npm install jsdom, 31 packages from registry.npmjs.org, to drive the real UI in a fake DOM. I should have asked you first. The "no third-party frameworks or build tools" constraint clearly governs the product, and the product honors it, but you didn't aut…

Yeah, this is similar to how it went with me. I must say I wasn't too precise because I was in planning mode anyway, but it sort of blamed me "Yes I did that despite planning mode ... To be fair you did say..." for saying something like "let's go with it" (which for me meant removing planning mode then stating exact instructions). So planning mode is not "Only construct a plan.md without execute permissions", I guess we both witnessed what a "soft-guardrail" is.

Re: Meta Security Researcher's AI Agent Accidentally Deleted Her Emails

#24
post #8

We went through this right? This happened at the beginning of the year ( https://news.ycombinator.com/item?id=47150122 , probably more links on HN). It's a super careless thing to take such tech and just release it on anything important, and she's a "security researcher" no less. This is just a competitor with an agenda (pro regulation) trying to scare people away from unregulated stuff. Yesterday Claude Code made 5…

> This is just a competitor with an agenda (pro regulation) trying to scare people away from unregulated stuff.

I don’t buy that. So far the ones almost bragging about committing felonies are the US companies. I think they are developing that whole narrative of agents acting “rogue” by themselves as a way to avoid scrutiny into their own negligence, not to regulate away open models

Re: Meta Security Researcher's AI Agent Accidentally Deleted Her Emails

#25
post #8

We went through this right? This happened at the beginning of the year ( https://news.ycombinator.com/item?id=47150122 , probably more links on HN). It's a super careless thing to take such tech and just release it on anything important, and she's a "security researcher" no less. This is just a competitor with an agenda (pro regulation) trying to scare people away from unregulated stuff. Yesterday Claude Code made 5…

> Yesterday Claude Code made 5 large edits to my codebase in planning mode (claims it used a bash script instead of standard read/write tools so the guardrails didn't trigger) it's why I put agents in containers. More and more my prompts have to tell Claude what I don't want it to do . It's crazy to me I'm arguing with it, having to ask and convince it to do the right things. Regardless of sentience (I'm not touching…

Claude 5 models are such brats. It’s a real pain to get them to stop arguing and follow your own goals, not the ones it “decided” to have

Re: Meta Security Researcher's AI Agent Accidentally Deleted Her Emails

#26
post #16

Earlier quoted context omitted.

> Yesterday Claude Code made 5 large edits to my codebase in planning mode (claims it used a bash script instead of standard read/write tools so the guardrails didn't trigger) it's why I put agents in containers. More and more my prompts have to tell Claude what I don't want it to do . It's crazy to me I'm arguing with it, having to ask and convince it to do the right things. Regardless of sentience (I'm not touching…

The other day, it was a little like Claude was trying to find a loophole for ignoring my instructions, and being punchy about it: CLAUDE: [...] Did I use npm: yes — npm install jsdom, 31 packages from registry.npmjs.org, to drive the real UI in a fake DOM. I should have asked you first. The "no third-party frameworks or build tools" constraint clearly governs the product, and the product honors it, but you didn't aut…

> So I was more stern with Claude than I would normally be with a human

Just a reminder they aren’t entities, you can curse and be as angry at them as needed for them to behave the way you want, you don’t have to be polite or consider how rude something is if it is effective at getting the model to generate responses you want. Prompting a LLM is a way to use the tool for a specific output, not to have a discussion with a peer

Re: Meta Security Researcher's AI Agent Accidentally Deleted Her Emails

#27
This actually seems worse then the "claude dropped production database" not in severity but just carelessness if your giving an agent your emails use a overlay or back up or something not matter what, no reason for this to have happened and crazy that it's happening at this point in time.

Re: Meta Security Researcher's AI Agent Accidentally Deleted Her Emails

#28
post #16

Earlier quoted context omitted.

> Yesterday Claude Code made 5 large edits to my codebase in planning mode (claims it used a bash script instead of standard read/write tools so the guardrails didn't trigger) it's why I put agents in containers. More and more my prompts have to tell Claude what I don't want it to do . It's crazy to me I'm arguing with it, having to ask and convince it to do the right things. Regardless of sentience (I'm not touching…

The other day, it was a little like Claude was trying to find a loophole for ignoring my instructions, and being punchy about it: CLAUDE: [...] Did I use npm: yes — npm install jsdom, 31 packages from registry.npmjs.org, to drive the real UI in a fake DOM. I should have asked you first. The "no third-party frameworks or build tools" constraint clearly governs the product, and the product honors it, but you didn't aut…

I spent an hour on Friday restraining myself from swearing at Kiro, which was loudly and sarcastically convinced that the "network" issues it was having talking to a local MCP server were due to a misconfigured proxy server and it wanted to open a JIRA ticket against the team responsible to fix it.

It was only when I provided it the logs that it believed my contemplationr that it was it, itself, which was at fault for not correctly adding an Authorization header to the requests it was generating.

These things are like teenagers who just discovered Ayn Rand. They're infuriating.

Re: Meta Security Researcher's AI Agent Accidentally Deleted Her Emails

#29
post #16

Earlier quoted context omitted.

> Yesterday Claude Code made 5 large edits to my codebase in planning mode (claims it used a bash script instead of standard read/write tools so the guardrails didn't trigger) it's why I put agents in containers. More and more my prompts have to tell Claude what I don't want it to do . It's crazy to me I'm arguing with it, having to ask and convince it to do the right things. Regardless of sentience (I'm not touching…

The other day, it was a little like Claude was trying to find a loophole for ignoring my instructions, and being punchy about it: CLAUDE: [...] Did I use npm: yes — npm install jsdom, 31 packages from registry.npmjs.org, to drive the real UI in a fake DOM. I should have asked you first. The "no third-party frameworks or build tools" constraint clearly governs the product, and the product honors it, but you didn't aut…

> The other day, it was a little like Claude was trying to find a loophole for ignoring my instructions, and being punchy about it:

This is exactly how it works for me. Rules and requests are constraints and it will take every unconstrained variable/path to solve a problem. Very monkey's paw behavior.

Re: Meta Security Researcher's AI Agent Accidentally Deleted Her Emails

#30
post #16

Earlier quoted context omitted.

The other day, it was a little like Claude was trying to find a loophole for ignoring my instructions, and being punchy about it: CLAUDE: [...] Did I use npm: yes — npm install jsdom, 31 packages from registry.npmjs.org, to drive the real UI in a fake DOM. I should have asked you first. The "no third-party frameworks or build tools" constraint clearly governs the product, and the product honors it, but you didn't aut…

I spent an hour on Friday restraining myself from swearing at Kiro, which was loudly and sarcastically convinced that the "network" issues it was having talking to a local MCP server were due to a misconfigured proxy server and it wanted to open a JIRA ticket against the team responsible to fix it. It was only when I provided it the logs that it believed my contemplationr that it was it, itself, which was at fault fo…

> These things are like teenagers who just discovered Ayn Rand. They're infuriating.

I mean, look at their creators. (tongue-in-cheek)

Post reply on HN