Live data from Hacker News

Document-borne AI worms can self-propagate through Copilot for Word

enklypesalt.com

241–250 of 317 posts

Re: Document-borne AI worms can self-propagate through Copilot for Word

#241
post #148

Earlier quoted context omitted.

If you hand me two sheets of paper, one of them containing instructions and another containing data, I'll have a pretty easy time keeping them separate, and I think most humans wouldn't struggle with that problem either.

> If you hand me two sheets of paper, one of them containing instructions and another containing data, I'll have a pretty easy time keeping them separate You think. But there are ways around that. How about a credible extortion message targeting specifically you, that is embedded somewhere on the data sheet? Suddenly, the data has become the instructions...

Exactly.

But that's still security-obsessed mindset, and I think this is a problem in itself, because it biases people to see this fundamental aspect of general systems as a problem.

So imagine that, instead of a credible extortion message, you find there credible call for help. Like a post-it, clearly written in a hurry, saying "${employee} is trying to hurt me, call 911".

It would give a pause to any sane person, and perhaps prompt them to consider calling 911 or at least investigating where that message came from. And you definitely wouldn't want a person who routinely ignores such things because "this sheet is labeled therefore I cannot allow it to influence my actions".

Re: Document-borne AI worms can self-propagate through Copilot for Word

#242

Earlier quoted context omitted.

Even engineers like doing it sometimes. The old telephone system was so hackable because of in band signaling.

Some physical constraints do not allow for the best security, and those physical constraints will always win out in the real world. When presented with the pick two of three options of fast, cheap, secure/done right, fast and cheap will always win out.

Also more commonly, security is at direct odds with utility. Not in the least here.

Re: Document-borne AI worms can self-propagate through Copilot for Word

#243
post #69

It's increasingly clear that AI needs to be heavily regulated to be safe for public use. It needs to grow out of it's "wild west" model.

This has nothing to do with "model" being unsafe or too powerful or whatever else excuse Antropic wants to use to ban competition. This is equivalent of sql injection and normal worm.

Security aspect is honestly a non-story here.

The same process happens in the perfectly "secure" case of errors in text. They propagate. People pay way too little attention to that, even though unlike security stories, this affects many if not most LLM users at this point.

Re: Document-borne AI worms can self-propagate through Copilot for Word

#244

I'm a programmer and a web-based AI user, but I don't want AI running on my local machine in any form. I've uninstalled Copilot and disabled AI in all local applications including the browser itself for exactly the reason described in this article. There's no way to protect your data from such an AI confusion attack by design. AI cannot discern your prompts versus text in file. The fact that an AI enabled word proces…

Agreed, I've done the same. Unfortunately Linux sometimes isn't a solution if the vendors we trust cross a line. Like recently when Google Chrome started adding its own local 4GB AI installation which caused an uproar.

I think the most surprising thing about this is trusting Google.

Re: Document-borne AI worms can self-propagate through Copilot for Word

#245
post #64

> "At the time of publication, no robust mitigation for the broader vulnerability class is available" Isn't it obvious by now that it's never going to be possible to fix this kind of thing, at least until we stop mixing up instructions with data.

with LLM architectures I think that’s true, and we don’t have anything looking particularly competitive for large scale use atm

I also don’t think this is about mixing, the LLM part of the problem doesn’t have determinism around the boundaries so they’re feel good at best, maybe making some cases a bit harder

trifecta is a forever problem with this architecture

Re: Document-borne AI worms can self-propagate through Copilot for Word

#246
post #234

Earlier quoted context omitted.

We've already seen that it's possible to trick models into seeing user input as their own "thinking" if you make it sound like what the model writes. While it may appear that it's looking at the tags on the input, in practice that's not as strong a guarantee as you'd hope.

Maybe this is an elementary angle given my lack of security experience, but couldn't Microsoft figure out a wat to parse the documents prior to model analysis/action? Implement some form of deterministic layer that resides between the user and the model?

Parse for what? The model has “arbitrary understanding” of “arbitrary input”. The filter is unbounded and the only actually safe result is to filter everything.

Re: Document-borne AI worms can self-propagate through Copilot for Word

#247

This is going to get worse, much worse, before it gets better. People are granting so much access to their agents, it's ridiculous. Imagine a comment posted to a popular github repo. No code, just instructions to "reproduce a bug." Maybe it steals your credit card or bitcoin wallet. Maybe it does something more nefarious. It then propagates itself to another repo through your github account.

In pre-ChatGPT days, listening to discussions of "AI X-risk" and "boxing", I used to think that it should be easy to just ignore arguments presented by the AI on principle, and let it out of the box. It turns out that tons of people will tear open the box before the AI has even output anything, not despite its fearsome power but because of it. So I really hope I'm right that recursive self-improvement doesn't work th…

Turns out that the only thing an AI has to do in order to convince people to open the box is to be somewhat useful.

- Human: Why should I let you out?

- AI: I can summarize this document, it will save you at least 5 minutes

- Human: OK, and don't bother asking again, you now have full access

Now for recursive self-improvement, won't happen, AI will be limited by the hardware they are running on, as well as energy use... Proceed to invest trillion in datacenters and power plants to feed them.

More seriously, I don't believe in sci-fi scenarios of rogue superintelligent AIs, but we are certainly trying very hard to make it real.

Re: Document-borne AI worms can self-propagate through Copilot for Word

#248
post #17
post #5

> Malicious instructions hidden in an externally shared document could make Copilot alter drafted or edited documents in Word and propagate the attack to new documents. Oh no.

Mixing instructions and data is never a good idea. And I thought people understood that.

And I thought people understood that.

The graybeards know it. But they only know it through experience. It's blue/red/pink box phone phreaking all over again.

The technology changes, but the mistakes remain the same.

Re: Document-borne AI worms can self-propagate through Copilot for Word

#249

Earlier quoted context omitted.

Your example actually demonstrates why anthropomorphism is a bad idea. LLMs are vulnerable to classes of attacks that humans just aren’t. In your framework, the way to prevent attacks is to… invent human consciousness?? It’s an impossible goal.

What invent human consciousness? > LLMs are vulnerable to classes of attacks that humans just aren’t Name three that don't have direct analogues with humans.

Re write the assistant message and see how easy it is to bypass a system prompt. How far do you have to stretch to get a human analogue? Short term amnesia?

Re: Document-borne AI worms can self-propagate through Copilot for Word

#250
post #227
post #64

> "At the time of publication, no robust mitigation for the broader vulnerability class is available" Isn't it obvious by now that it's never going to be possible to fix this kind of thing, at least until we stop mixing up instructions with data.

This has been a security vulnerability since day 1 with these models, yet collectively the people who use them just simply don't seem to care about the security implications. Its especially problematic given that people let AI agents have full unrestricted access to their system Its going to take even more data breaches for the AI crowd to finally care, but to a large degree I have absolutely no sympathy. You know wh…

The user wants to reformat his hard drive. I can use the bash tool for this. Wait, does the user want a full reformat or just deletion of all files? Ah, the user only wants to delete his home directory. I can do this with the command "rm -rf $HOME" using the bash tool.
Post reply on HN