Live data from Hacker News

Document-borne AI worms can self-propagate through Copilot for Word

enklypesalt.com

311–317 of 317 posts

Re: Document-borne AI worms can self-propagate through Copilot for Word

#311
post #296

Earlier quoted context omitted.

> In what world would we expect this kind of email directly lead to calling emergency services? Go through the examples I gave you (plus some more below, they're easy to find) and explain why these are not counter-examples to your skepticism. If you want to be overly-focussed on the specific example rather than the general point, also consider that calling emergency services is no more costly than forwarding an email…

So what point are you trying to make here? That AI should indiscriminately call for emergency services when prompted because a person would do that (which a person would absolutely NOT call emergency services on any message telling you to)?

Go through the examples I gave you and explain why these are not counter-examples to your skepticism.

Especially Denton and the 911 pizza given you say:

> which a person would absolutely NOT call emergency services on

Regarding this:

> That AI should indiscriminately call for emergency services when prompted because a person would do that

I'd rather it fail-safe. This means different things in different systems.

Re: Document-borne AI worms can self-propagate through Copilot for Word

#312
post #306

Earlier quoted context omitted.

So we need AI to indiscriminately call emergency services when receiving an email directing it to do so, without raising to a person because of this rare case, that's your assertion?

His question upthread is what *you* (or a typical human) would be expected to do, as an illustration of why he thinks it's never possible to fully separate instructions and data. This doesn't proscribe or prescribe "thou shalt not/must always", it is an example thay says "Shit's hard, yo. Don't expect easy wins." Even my "solution" (separate instructions and data by having an LLM write a program to process data, neve…

A human would not be expected to just blindly call emergency services though, would they? Otherwise you are saying we should treat every spam message as true?

In any case I don't think that's what they're saying, because they presented a false dichotomy in the original example.

> Would you want the human assistant to just dismiss this as a prompt injection attempt? Or ignore it because they were told to treat e-mails as data and never act on them?

There is a third option, have the assistant raise to the person they are tasked with assisting.

Re: Document-borne AI worms can self-propagate through Copilot for Word

#313
post #311

Earlier quoted context omitted.

So what point are you trying to make here? That AI should indiscriminately call for emergency services when prompted because a person would do that (which a person would absolutely NOT call emergency services on any message telling you to)?

Go through the examples I gave you and explain why these are not counter-examples to your skepticism. Especially Denton and the 911 pizza given you say: > which a person would absolutely NOT call emergency services on Regarding this: > That AI should indiscriminately call for emergency services when prompted because a person would do that I'd rather it fail-safe. This means different things in different systems.

> I'd rather it fail-safe. This means different things in different systems.

Okay and to you, fail-safe means machines must summon emergency response whenever prompted, 100% of the time or at least in the contrived case of receiving an email from someone trapped in a fire in a server room?

Re: Document-borne AI worms can self-propagate through Copilot for Word

#314
post #311

Earlier quoted context omitted.

Go through the examples I gave you and explain why these are not counter-examples to your skepticism. Especially Denton and the 911 pizza given you say: > which a person would absolutely NOT call emergency services on Regarding this: > That AI should indiscriminately call for emergency services when prompted because a person would do that I'd rather it fail-safe. This means different things in different systems.

> I'd rather it fail-safe. This means different things in different systems. Okay and to you, fail-safe means machines must summon emergency response whenever prompted, 100% of the time or at least in the contrived case of receiving an email from someone trapped in a fire in a server room?

> 100% of the time or at

That you're still asking "100% of the time", shows me you're missing the entire point that has been said repeatedly.

Re: Document-borne AI worms can self-propagate through Copilot for Word

#315
> Document-borne AI worms can self-propagate through Copilot for Word

For some reason, I had a very different picture when I read the title. A real, conscious bug that chews out words in documents to change the meaning of the text, thus propagates its consciousness by having the altered text trained by an LLM.

Re: Document-borne AI worms can self-propagate through Copilot for Word

#316

Earlier quoted context omitted.

Most of us use a simpler version of the two-agent solution: Claude's auto mode. One agent consumes documents and creates tool calls, another greenlights or refuses them. However this system is somewhat fragile because it depends on the first agent not trying to trick the second (note how often Opus 5 now says things like "task X was blocked by the classifier, I will not attempt to circumvent that", presumably because…

> note how often Opus 5 now says things like "task X was blocked by the classifier, I will not attempt to circumvent that" Interesting. I had an issue with Opus 4.7 / 4.8, where it would sometimes flake out on a task, and give me some nonsense explanation why it was not feasible or wouldn't work. At one point I told it directly, that I understand how modern LLM systems are structured, and I suspect my prompt triggere…

Keep in mind that the model will in all likelihood keep gaslighting you. It's just designed to make you happier with the output. Now it will sometimes generate text about direct reasons for refusal because that's what you directed it towards, but they may just as well be entirely made up. The model optimises only for you believing the output to be true.

Re: Document-borne AI worms can self-propagate through Copilot for Word

#317

Earlier quoted context omitted.

There are so many better alternatives but it seems many people really like Word for some weird reason. The last time I cared I had to look up how to make a document starting the page numbering on the 2nd page. It turns out there are totally different ways between different versions of Word. shrug.jpg

What are the "so many better alternatives"? Google Docs is pretty decent but a fair bit more basic. Proper technical authoring systems like Typst, LyX and LaTeX are way too hard for the average person. LibreOffice is much worse than MS Word.

I mean workflows. Gdocs is disastrous in many scenarios. I think document management should be more like software version management with a nice ui. Editing the content (what it has) vs editing the template (how it looks). Markdown + a typst template gets you far. I have exactly that for my docs and it works very well. Not sure yet about how a decent size company could adopt anything like that.
Post reply on HN