Live data from Hacker News

What happened after 2k people tried to hack my AI assistant

fernandoi.cl

181–186 of 186 posts

Re: What happened after 2k people tried to hack my AI assistant

#185
post #41

Earlier quoted context omitted.

Yeah, I remember some ad by an LLM security company hitting HN a year or so with a "challenge" to do prompt injection. The final level was their product and it was impossible. But it was also impossible to get the LLm to do _anything_. May as well just echo "prompt injection attempt detected" at that point and never send anything to an LLM.

This one? https://gandalf.lakera.ai/baseline I remember doing it and getting quite far, but not completely beating it. I know some other people did beat it completely though.

We're building open source challenges where you can inspect the actual agent to see if it's possible or not. We're planning to revamp this over the next couple of days and maintain a weekly cadence. https://playground.fabraix.com/
Post reply on HN