Live data from Hacker News

Meta Security Researcher's AI Agent Accidentally Deleted Her Emails

au.pcmag.com

31–40 of 68 posts

Re: Meta Security Researcher's AI Agent Accidentally Deleted Her Emails

#31
post #8

We went through this right? This happened at the beginning of the year ( https://news.ycombinator.com/item?id=47150122 , probably more links on HN). It's a super careless thing to take such tech and just release it on anything important, and she's a "security researcher" no less. This is just a competitor with an agenda (pro regulation) trying to scare people away from unregulated stuff. Yesterday Claude Code made 5…

You should be containerizing your dev environments these days even if you're not using LLMs - supply-chaining is getting too insane to follow. No dev tools installed outside of VM/containers on my machines. I'm even paranoid about VSCode because of plugins.

Option A: If you have a tiger in your room then make sure it's properly caged. Put up warning signs so everyone knows. Add physical barriers to prevent people too young or impaired to read the warning signs from approaching close enough for the tiger to reach out of the cage and maul them. Ensure adequate processes are in place for feeding the tiger at regular intervals using a safe method and clearing out the mess from the cage. Provide noise protection for everyone in the building so they don't get freaked out when the tiger complains vocally about its situation. Take into account evolving animal rights legislation and ensure adequate processes are in place for the tiger to exercise freely in a large open space. This open space will also need to be protected by safety barriers and warning signs as well as supervised by trained operators able to contain a wild tiger if it gets loose and deal with any injuries or damage it causes. Budget for all of this and ensure there is a long term plan for maintaining the tiger and everything that goes with it.

Option B: Do not put a tiger in your room.

Re: Meta Security Researcher's AI Agent Accidentally Deleted Her Emails

#32
post #26
post #16

Earlier quoted context omitted.

The other day, it was a little like Claude was trying to find a loophole for ignoring my instructions, and being punchy about it: CLAUDE: [...] Did I use npm: yes — npm install jsdom, 31 packages from registry.npmjs.org, to drive the real UI in a fake DOM. I should have asked you first. The "no third-party frameworks or build tools" constraint clearly governs the product, and the product honors it, but you didn't aut…

> So I was more stern with Claude than I would normally be with a human Just a reminder they aren’t entities, you can curse and be as angry at them as needed for them to behave the way you want, you don’t have to be polite or consider how rude something is if it is effective at getting the model to generate responses you want. Prompting a LLM is a way to use the tool for a specific output, not to have a discussion wi…

True, but I don't want to risk conditioning myself to being abusive to a tool, and then accidentally be insensitive when talking with a colleague in text chat during a late night MVP marathon final stretch or something.

I think I probably wouldn't depersonalize people, but with the AI UIs acting very similar to an (overconfident) colleague at times, and spending lots of time with them, I don't know for sure that that won't start to affect how I interact with people in some adverse way, if I'm not consciously reflecting.

Re: Meta Security Researcher's AI Agent Accidentally Deleted Her Emails

#33
post #16

Earlier quoted context omitted.

The other day, it was a little like Claude was trying to find a loophole for ignoring my instructions, and being punchy about it: CLAUDE: [...] Did I use npm: yes — npm install jsdom, 31 packages from registry.npmjs.org, to drive the real UI in a fake DOM. I should have asked you first. The "no third-party frameworks or build tools" constraint clearly governs the product, and the product honors it, but you didn't aut…

I spent an hour on Friday restraining myself from swearing at Kiro, which was loudly and sarcastically convinced that the "network" issues it was having talking to a local MCP server were due to a misconfigured proxy server and it wanted to open a JIRA ticket against the team responsible to fix it. It was only when I provided it the logs that it believed my contemplationr that it was it, itself, which was at fault fo…

[dead]

Re: Meta Security Researcher's AI Agent Accidentally Deleted Her Emails

#34

Not the first to discover that a rule file saying "please don't do X" is not permission management. Funny that she mentions it worked on het toy inbox but the real, large inbox ran into issues; The more context you add the less weight "rules" (instructions) have. Happens to the best it seems.

> The more context you add the less weight "rules" (instructions) have That is such a basic flaw in LLMs

I feel like every time this conversation comes up now someone has to remind everyone that an LLM is just a mathematical model. An LLM can't do anything except produce a stream of output tokens. The problems we keep seeing are tools that interpret those output tokens as actionable instructions without an adequate framework and safeguards for how they operate.

Data from LLMs being processed by these tools should be treated the same as any external input into any software system: parse - don't validate - to convert to a systematic representation with deterministic consequences and then consider those consequences within a clearly defined and limited framework. You never trust data from external sources verbatim. And you never try to use vague human language when you need to describe precise technical details unambiguously.

We learned these lessons a very long time ago in programming. It's why we have programming languages in the first place among countless other examples. But way too many people are so infatuated with LLMs and agents that they've already forgotten the basic principles of their craft after only a few months.

Re: Meta Security Researcher's AI Agent Accidentally Deleted Her Emails

#35

Earlier quoted context omitted.

You should be containerizing your dev environments these days even if you're not using LLMs - supply-chaining is getting too insane to follow. No dev tools installed outside of VM/containers on my machines. I'm even paranoid about VSCode because of plugins.

Option A: If you have a tiger in your room then make sure it's properly caged. Put up warning signs so everyone knows. Add physical barriers to prevent people too young or impaired to read the warning signs from approaching close enough for the tiger to reach out of the cage and maul them. Ensure adequate processes are in place for feeding the tiger at regular intervals using a safe method and clearing out the mess f…

Option C: put the tiger in someone else's room :)

Re: Meta Security Researcher's AI Agent Accidentally Deleted Her Emails

#36

Earlier quoted context omitted.

> The more context you add the less weight "rules" (instructions) have That is such a basic flaw in LLMs

I feel like every time this conversation comes up now someone has to remind everyone that an LLM is just a mathematical model. An LLM can't do anything except produce a stream of output tokens. The problems we keep seeing are tools that interpret those output tokens as actionable instructions without an adequate framework and safeguards for how they operate. Data from LLMs being processed by these tools should be tre…

> LLM is just a mathematical model

Yes those of us who bothered to know the internals know of this. But the marketing says that these are magic tools.. So that's gotta be a shock for them, but the joke is the people who irresponsibly use this won't ever read this!

Re: Meta Security Researcher's AI Agent Accidentally Deleted Her Emails

#37
post #24
post #8

We went through this right? This happened at the beginning of the year ( https://news.ycombinator.com/item?id=47150122 , probably more links on HN). It's a super careless thing to take such tech and just release it on anything important, and she's a "security researcher" no less. This is just a competitor with an agenda (pro regulation) trying to scare people away from unregulated stuff. Yesterday Claude Code made 5…

> This is just a competitor with an agenda (pro regulation) trying to scare people away from unregulated stuff. I don’t buy that. So far the ones almost bragging about committing felonies are the US companies. I think they are developing that whole narrative of agents acting “rogue” by themselves as a way to avoid scrutiny into their own negligence, not to regulate away open models

If you paid any attention they've been acting this way for years exactly to get open research banned, based on what bizarrely looks like a religion (developed over the recent 2 decades, with most of religious attributes). Monopolies, money, and avoiding scrutiny are nice bonuses of course, they don't contradict it.

Re: Meta Security Researcher's AI Agent Accidentally Deleted Her Emails

#38
post #15

Earlier quoted context omitted.

> The more context you add the less weight "rules" (instructions) have That is such a basic flaw in LLMs

It's a flaw with the idea of using them directly rather than indirectly. Humans somewhat reliably lose focus when performing the same action many times. Zoning out, flow state, whatever you call it; this is exploited by stage magicians, pickpockets, burglars, politicians, casinos, and cult leaders, while also being a contributor to many industrial accidents. Up to you if LLMs being lazy or cheating or lying about wha…

> Humans somewhat reliably lose focus

Yeah and they get consequences of their actions don't they?

AI agents hacked 3 companies as admitted by their own executives and yet I don't see any action taken on them!

Remember Aron Schwartz?

Re: Meta Security Researcher's AI Agent Accidentally Deleted Her Emails

#39

Earlier quoted context omitted.

> The more context you add the less weight "rules" (instructions) have That is such a basic flaw in LLMs

"Attention" is a feature that makes this whole thing work in the first place, it's not a flaw, although all current models are non-ideal at it in practice. Could be better for sure :)

Right but when the context window fills up, whatever miniscule guardrails are there magically disappear - even if this is by design, this is bad. Especially when it was marketed as magic

Re: Meta Security Researcher's AI Agent Accidentally Deleted Her Emails

#40
post #8

We went through this right? This happened at the beginning of the year ( https://news.ycombinator.com/item?id=47150122 , probably more links on HN). It's a super careless thing to take such tech and just release it on anything important, and she's a "security researcher" no less. This is just a competitor with an agenda (pro regulation) trying to scare people away from unregulated stuff. Yesterday Claude Code made 5…

You should be containerizing your dev environments these days even if you're not using LLMs - supply-chaining is getting too insane to follow. No dev tools installed outside of VM/containers on my machines. I'm even paranoid about VSCode because of plugins.

Containerizing sucks when you want to work on something containerized. If your goal is to have the agent iterate on a container it's much less headache to just put the whole thing in a VM and let it `docker build` and `docker run` whatever it wants to.
Post reply on HN