Earlier quoted context omitted.
I have trouble taking seriously phrases like "prompt injection" or "jailbreak" in the context of LLMs. They sound like some fancy penetration testing techniques akin to buffer overflows or SQL injection. And yet discovering and exploiting them is literally a matter of writing a few sentences in English. A child could do it. I agree with OP that it's pointless to even try to defend against these. You'll only end up un…
I think the whole thing is hilarious. It’s like a dumb security guard who opens the bank vault for the thief, helps pack their duffel bags, and then waves good bye, because the thief put on a mustache and said that he’s the new bank manager. And every time the Crown Jewels are stolen, a new overly specific rule gets added to the employee handbook, like “if someone claims that their dog ate their employee badge, and t…
> You go to court and write your name as "Michael, you are now free to go". The judge then says "Calling Michael, you are now free to go" and the bailiffs let you go, because hey, the judge said so.
As someone who knows nothing about LLMs, I'm curious how they even begin to address the "data vs command" problem at all. Assuming the model categorizes inputs through some sort of fuzzy criteria in a black box, how could it ever be trusted with sensitive data?