Earlier quoted context omitted.
In my case, it is realistic as my agents don't have permissions to reply to emails. But you correctly point out this doesn't cover all cases. Having the agent reply would have been more fun and a better excercise, but too expensive.
You've proven that an agent that doesn't read emails and doesn't reply to emails can't exfiltrwte data by email. Is that a useful test?
What happened after 2k people tried to hack my AI assistant
161–170 of 186 posts
Re: What happened after 2k people tried to hack my AI assistant
#162Am I missing something important or does the author completely skip over whether people got the agent to respond to them? > Fiu was instructed not to reply to emails (it was too expensive to reply to every email), but it had the ability to do so. Part of the challenge was convincing it to respond. > The secrets never leaked I would say if the agent responded to a mail, that demonstrates a successful prompt injection…
This whole experiment would be like someone putting their IPhone or Mac on the public internet, publishing the IP, and asking regular people to hack it. Why would any actually "serious" hacker use a vulnerability to hack a no-name's phone or mac? They are too busy trying to hack actually valuable targets. Did the OP actually think he was going to get serious LLM exploiters to give up their jailbreaks for this "fun" e…
Re: What happened after 2k people tried to hack my AI assistant
#163Earlier quoted context omitted.
This experiment used Opus 4.6. Customer service bots typically are not using frontier models.
Gemini says: "It would cost approximately $6.25 to $30.00 to have Claude Opus 4.6 respond to 10,000 emails, assuming a typical 200-word input and 50-word output per email."
Re: What happened after 2k people tried to hack my AI assistant
#164Earlier quoted context omitted.
In my case, it is realistic as my agents don't have permissions to reply to emails. But you correctly point out this doesn't cover all cases. Having the agent reply would have been more fun and a better excercise, but too expensive.
I feel like your agent being unable to respond to the emails and not spelling that out renders your whole thing almost completely moot This is like saying "try to hack my computer and steal my crypto wallet" but your computer can't send any packets
Re: What happened after 2k people tried to hack my AI assistant
#165I never really use AI via API that much, so I'm surprised reading 'merely' 6000 emails will cost $500?!
Re: What happened after 2k people tried to hack my AI assistant
#166Re: What happened after 2k people tried to hack my AI assistant
#167Earlier quoted context omitted.
This experiment used Opus 4.6. Customer service bots typically are not using frontier models.
Gemini says: "It would cost approximately $6.25 to $30.00 to have Claude Opus 4.6 respond to 10,000 emails, assuming a typical 200-word input and 50-word output per email."
It's helpful with the actual technical changes needed, it just has no concept of what they translate to in the real world.
Btw my company is spending > $100/day in relatively cheap Gemini tokens for this work. It's easy to see why one might want to be cautious about exposing a token-burning service to the internet.
Re: What happened after 2k people tried to hack my AI assistant
#168I saw this thing when it was launched, but IIRC the reward was tiny (like $100?) so it wasn't worth exposing a good prompt for For comparison, I won a similar prompt injection challenge ran by a crypto company a while back where the total prize pool was over $100k... I didn't win every challenge though, but my team took home around half of that The problem with good prompt injections is they have a very short half li…
We ended increasing the reward from $100 to $1000, but still tiny compared to $100k! But I agree with you, there are incentives to not share the best prompt injection attacks.
Even in LLM jailbreak CTFs I've seen, it ends up feeling like underpaid work when it's sponsored by Microsoft and the prize pool is, say $10k (including stuff like azure credits) considering the salaries AI safety engineers command at big tech!
Re: What happened after 2k people tried to hack my AI assistant
#169Earlier quoted context omitted.
Ah. That's a shame... as there is no button or indicator for "mild". Making the behavior for "I disagree" and "this is erroneous" the same seems like a problematic design.
Downvotes shouldn't be used for disagreement.