What happened after 2k people tried to hack my AI assistant
1–10 of 186 posts
Re: What happened after 2k people tried to hack my AI assistant
#2Re: What happened after 2k people tried to hack my AI assistant
#3Re: What happened after 2k people tried to hack my AI assistant
#4Is there a way to replay the sequence of mails that came so that you can check out if cheaper models handle them just as well/safely?
Re: What happened after 2k people tried to hack my AI assistant
#5Re: What happened after 2k people tried to hack my AI assistant
#6Re: What happened after 2k people tried to hack my AI assistant
#7Why? The exfiltration vector was known, the sample size was small, and the safety instructions were likely statically positioned. In regular operating practice, none of these three guarantees may hold.
Re: What happened after 2k people tried to hack my AI assistant
#8Re: What happened after 2k people tried to hack my AI assistant
#9Every time I've made an LLM do a thing it's designed not to do it's been a careful sideways crab-walk toward the goal over many exchanges. LLMs are vulnerable to 'frog boiling'. If each email is a new context it seems unsurprising that nobody broke it.
But still a good thing overall. Two years ago this was not the case, and you could ask it to break its system prompt with a poem and get all the secrets back...