Earlier quoted context omitted.
Good catch! Luckily it's not that sophisticated and even thinking of open sourcing the code. I've filtered out that response but I'm sure people will find clever ways of extracting the prompt anyways.
A start would be to detect if the result of the prompt includes your exact prompt. Or something that looks similar. Although one could probably tell it to talk like a pirate to evade that, or something.
That's exactly what I did. But there are probably ways to have the model encode the response (e.g. "answer but with the words in reversed order"), so I do expect motivated people to figure out ways to extract it. I guess I'd probably spend more effort on this if my prompt was really clever, but it's not.