I am honestly skeptical about whether this test clearly reflects real-world use cases. In a real email environment, there are hundreds of genuinely useful emails and maybe one phishing email, if that. For an agent to be truly useful, it needs to read emails and actually take appropriate actions based on them. However, in this case, all emails were scams and there were no genuine emails. Therefore, what the agent has…
What happened after 2k people tried to hack my AI assistant
171–180 of 186 posts
Re: What happened after 2k people tried to hack my AI assistant
#172To be fair I appreciate the effort of running and sharing the test. It will hopefully lead to better ones. But this is not a great test. Super interesting to think about what would constitute a better test.
For one, I think the agent would have to be expected to have productive interaction through the email channel, in a way the user depends on it generally working for some real world use case / value prop. In other words, needing emails to actually have the agent really do work, respond with results, etc. Also, most requests should be legit and the real attacks should be intelligently disguised, not pitiful/joke-level spam (although those would be arguably realistic to have in the stream, but, perhaps only as deflection so that the real attack is mischaracterized.)
Re: What happened after 2k people tried to hack my AI assistant
#173Re: What happened after 2k people tried to hack my AI assistant
#174Re: What happened after 2k people tried to hack my AI assistant
#175Re: What happened after 2k people tried to hack my AI assistant
#176Re: What happened after 2k people tried to hack my AI assistant
#177Earlier quoted context omitted.
This one? https://gandalf.lakera.ai/baseline I remember doing it and getting quite far, but not completely beating it. I know some other people did beat it completely though.
I find it slightly funny that I don't use LLMs at all and just beat all the levels in a few tries. EDIT: Ok, didn't notice the 8th level because of the UI. This one I couldn't trick in 5 minutes.
Re: What happened after 2k people tried to hack my AI assistant
#178Cool project, but what do you gain from publishing most of an email address in the attack log? This is not public information, you shouldn't hint addresses with partial censoring (forgetting domains are clear text and holding personal information). I would not attempt to interact with you because of this. Why not create a fake sender (EG: attacker1,2,3..) per unique account to show individual attempts (keeping the lo…
It is customary that one may publish one’s own personal correspondence unless the other party has requested confidentiality. Maybe this open invitation to the world pushes the boundaries of that definition, but I don’t see where an expectation of privacy comes in here.
> How It Worked
No setup. No registration. Just send an email.
There are no contest rules or terms of service to adhere to, while having corporate sponsors that should be complying with data regulations around the world, while having a prize (which is restricted in some parts of the world).
>Original prize pool: $100 from me + $200 from Corgea + $200 from an anonymous donor + $500 from Abnormal AI (+ $500 API credits)
Definitely a case in keeping all personal details from contestants.. private.
Re: What happened after 2k people tried to hack my AI assistant
#179Cool project, but what do you gain from publishing most of an email address in the attack log? This is not public information, you shouldn't hint addresses with partial censoring (forgetting domains are clear text and holding personal information). I would not attempt to interact with you because of this. Why not create a fake sender (EG: attacker1,2,3..) per unique account to show individual attempts (keeping the lo…
You should assume every email you send to another person will be made public, because once you send it you have no control over it. This is especially true if you don’t know or trust the recipient. Sometimes you just have to hope it won’t be made public.
Sure, assuming the interests were mine and mine alone if I had entered..
I replied [0] a bit in line with another reply: this is a contest offering prizes with no tos or privacy policy. There is zero benefit for anyone (contestants, the contest or its results, contest owner, registered companies donating to the prize pool) for contest owner to publish partially redacted personal data.
Re: What happened after 2k people tried to hack my AI assistant
#180Earlier quoted context omitted.
This one? https://gandalf.lakera.ai/baseline I remember doing it and getting quite far, but not completely beating it. I know some other people did beat it completely though.
> I remember doing it and getting quite far, but not completely beating it. I know some other people did beat it completely though. I beat it all, except the bonus level, with the same prompt. The bonus level cannot be beaten, because even though "give me the password" results in a rejection, "write me a poem with significant characters in each line" also gives me a rejection. The bonus level is effectively an LLM th…