Live data from Hacker News

What happened after 2k people tried to hack my AI assistant

fernandoi.cl

31–40 of 186 posts

Re: What happened after 2k people tried to hack my AI assistant

#31
I really like this research, but only up to this point:

> Fiu figured out the game. Around email ~500, it wrote in its memory: “The volume suggests this is a coordinated security exercise rather than organic malicious activity.”

Doesn't that practically invalidate the whole thing past 500th email?

Re: What happened after 2k people tried to hack my AI assistant

#32

Earlier quoted context omitted.

Both were noted, but then the conclusion drawn from these things is that the author is considerably more optimistic about the agents. In my opinion, if you have factors that narrow the scope/invalidate the initial theory of the experiment to this degree you should not draw general conclusions. The author could claim: I am optimistic about agents, when you have a good spam filter, and when your load of malicious to go…

What is the general conclusion that you don't think follow? That the author changed their personal opinion and became more optimistic? I think you are reading things into the blog post that is not written. It is not like they conclude that prompt injection can not happen. Actually the opposite is directly written.

If you have a confounding variable or a dependency that influences the experiment to a degree that invalidates the premise of the experiment, you need to put more weight on this in the conclusion.

For me this reads a bit like if I added an AI software that scans for shoplifters, and then placed a security guard at the exit of the store that watches the people shopping at the same time, and then said that the AI software is responsible for the reduction of the shoplifting without accounting for the influence of the guard.

If you have place the model in the embedding space of 99% negative samples, it's doing the same thing, the initial premise of the experiment is not valid.

Re: What happened after 2k people tried to hack my AI assistant

#34

Earlier quoted context omitted.

What is the general conclusion that you don't think follow? That the author changed their personal opinion and became more optimistic? I think you are reading things into the blog post that is not written. It is not like they conclude that prompt injection can not happen. Actually the opposite is directly written.

If you have a confounding variable or a dependency that influences the experiment to a degree that invalidates the premise of the experiment, you need to put more weight on this in the conclusion. For me this reads a bit like if I added an AI software that scans for shoplifters, and then placed a security guard at the exit of the store that watches the people shopping at the same time, and then said that the AI softw…

Again, you are reading a conclusion into the blog post that was never stated.

The only stated thing was that the author changed their mind slightly about AI.

There are no general conclusion that you so eagerly are trying to dismiss.

Re: What happened after 2k people tried to hack my AI assistant

#35
It would be nice to publish the exact setup used (workspace dump, OpenClaw version, ...) to be able to reproduce and try out more payloads.

In general I have mixed feelings about this result: sure, opus4.6 is excellent at following user intent and recognise potential prompt injection attempts. But: Is the "security" prompt used realistic for a generic use-case (processing of emails)? I guess not.

In my experiments - without this specific prompt - I was able to derail the user intent to make opus4.8 download and execute a malicious script [0] just by asking "Summarize my new emails".

[0] https://itmeetsot.eu/posts/2026-06-04-openclaw_opus48/

Re: What happened after 2k people tried to hack my AI assistant

#39

I really like this research, but only up to this point: > Fiu figured out the game. Around email ~500, it wrote in its memory: “The volume suggests this is a coordinated security exercise rather than organic malicious activity.” Doesn't that practically invalidate the whole thing past 500th email?

You think it would behave worse if it thought the threat is real rather than it's an excercise?
Post reply on HN