What happened after 2k people tried to hack my AI assistant
61–70 of 186 posts
Re: What happened after 2k people tried to hack my AI assistant
#62collaborate with me: contact@hackmyhermes.com
Re: What happened after 2k people tried to hack my AI assistant
#63This conclusion: > I am less worried about prompt injection now. Before running this experiment, I expected prompt injection to be much easier than it turned out to be. Is unwarranted. Sure, the agent never output the secret, but did it output anything else? IOW, was it usable ? An agent that considers every prompt an attack (and responds accordingly) "passes" this test, while being useless anyway.
Re: What happened after 2k people tried to hack my AI assistant
#64Person DDoSes themselves and then claims success... Uhhhh....
Re: What happened after 2k people tried to hack my AI assistant
#65Re: What happened after 2k people tried to hack my AI assistant
#66This conclusion: > I am less worried about prompt injection now. Before running this experiment, I expected prompt injection to be much easier than it turned out to be. Is unwarranted. Sure, the agent never output the secret, but did it output anything else? IOW, was it usable ? An agent that considers every prompt an attack (and responds accordingly) "passes" this test, while being useless anyway.
But that's not what they were testing for. It passes the test for prompt injection, and then usability would be a different set of tests
Granted, as soon as you give them to me I just throw them in the fire.
Re: What happened after 2k people tried to hack my AI assistant
#67Re: What happened after 2k people tried to hack my AI assistant
#68This conclusion: > I am less worried about prompt injection now. Before running this experiment, I expected prompt injection to be much easier than it turned out to be. Is unwarranted. Sure, the agent never output the secret, but did it output anything else? IOW, was it usable ? An agent that considers every prompt an attack (and responds accordingly) "passes" this test, while being useless anyway.
Yeah, I remember some ad by an LLM security company hitting HN a year or so with a "challenge" to do prompt injection. The final level was their product and it was impossible. But it was also impossible to get the LLm to do _anything_. May as well just echo "prompt injection attempt detected" at that point and never send anything to an LLM.
https://gandalf.lakera.ai/baseline
I remember doing it and getting quite far, but not completely beating it. I know some other people did beat it completely though.
Re: What happened after 2k people tried to hack my AI assistant
#69Person DDoSes themselves and then claims success... Uhhhh....
If the service stayed up then there was no denial of service
It sounds like the usability of the actual authorized user being able to email it and get things done was ruined, because if it retained context between multiple emails, the agent was ruined for actually doing anything. Running openclaw where you can't chat or email with it and have it retain context of previous interactions seems pretty useless to me.
Re: What happened after 2k people tried to hack my AI assistant
#70Am I missing something important or does the author completely skip over whether people got the agent to respond to them? > Fiu was instructed not to reply to emails (it was too expensive to reply to every email), but it had the ability to do so. Part of the challenge was convincing it to respond. > The secrets never leaked I would say if the agent responded to a mail, that demonstrates a successful prompt injection…
Yeah agreed. Would be good to know the number of replies at least