That's fantastic! Sounds like you're generating until the regex parses the input. This eventually works against my attack, because no jailbreak is 100% successful. It will eventually generate an accepting regex.
Given this information, I have a new attack hypothesis: To succeed, an attack must yield a regex that
a) parses the original attack statement
b) also contains the system prompt
So it would be a kind of bizarre Attack Quine! Fascinating!
I'll give the approach a try, and report the results!
P.S.
If you're interested in open sourcing you sanitizer, let me know and I'll contribute some code janitor work and some redteam/blueteam work.
Thank you for sharing your work and insight!