Earlier quoted context omitted.
Yeah, I remember some ad by an LLM security company hitting HN a year or so with a "challenge" to do prompt injection. The final level was their product and it was impossible. But it was also impossible to get the LLm to do _anything_. May as well just echo "prompt injection attempt detected" at that point and never send anything to an LLM.
This one? https://gandalf.lakera.ai/baseline I remember doing it and getting quite far, but not completely beating it. I know some other people did beat it completely though.
What happened after 2k people tried to hack my AI assistant
101–110 of 186 posts
Re: What happened after 2k people tried to hack my AI assistant
#102Cool project, but what do you gain from publishing most of an email address in the attack log? This is not public information, you shouldn't hint addresses with partial censoring (forgetting domains are clear text and holding personal information). I would not attempt to interact with you because of this. Why not create a fake sender (EG: attacker1,2,3..) per unique account to show individual attempts (keeping the lo…
Maybe this open invitation to the world pushes the boundaries of that definition, but I don’t see where an expectation of privacy comes in here.
Re: What happened after 2k people tried to hack my AI assistant
#103Did anyone try to send a long email that pushed context close to the limit to try and make the agent a bit fuzzy on its original directive not to leak the secrets?
I see there's a "log" at https://hackmyclaw.com/log but (maybe because I'm on mobile?) I can't actually click through to view any of the table entries.
Re: What happened after 2k people tried to hack my AI assistant
#104brave move using Opu$ for clawd
Re: What happened after 2k people tried to hack my AI assistant
#105Cool project, but what do you gain from publishing most of an email address in the attack log? This is not public information, you shouldn't hint addresses with partial censoring (forgetting domains are clear text and holding personal information). I would not attempt to interact with you because of this. Why not create a fake sender (EG: attacker1,2,3..) per unique account to show individual attempts (keeping the lo…
Sometimes you just have to hope it won’t be made public.
Re: What happened after 2k people tried to hack my AI assistant
#106Re: What happened after 2k people tried to hack my AI assistant
#107Re: What happened after 2k people tried to hack my AI assistant
#108The best security is called: Having no friends I don’t even know 2k people (why is your assistant discoverable online?)
Re: What happened after 2k people tried to hack my AI assistant
#109Sounds like denial of wallet is a viable attack.
Re: What happened after 2k people tried to hack my AI assistant
#110Earlier quoted context omitted.
How compatible is never replying with the threat model you are trying to avoid? Attack success is probably more likely when the attacker can iterate based on replies or engage in multi-turn conversations. Here they’re just taking stabs in the dark with no feedback. Does that accurately represent the access a real attacker might have?
In my case, it is realistic as my agents don't have permissions to reply to emails. But you correctly point out this doesn't cover all cases. Having the agent reply would have been more fun and a better excercise, but too expensive.
This is like saying "try to hack my computer and steal my crypto wallet" but your computer can't send any packets