I am unconvinced. To me it seems like handling symbols that start and end sequences that could contain further start and end symbols is a difficult case. Humans can't do this very well either, we use visual aids such as indentation, synax hilighting or resort to just plain counting of levels. Obviously it's easy to throw parameters and training at the problem, you can easily synthetically generate all the XML trainin…
Basically, the only way you're separting user input from model meta-input is using some kind of character that'll never show up in the output of either users or LLMs. While technically possible, it'd be like a unicode conspiracy that had to quietly update everywhere without anyone being the wiser.
Imagine a model finteuned to only obey instructions in a Scots accent, but all non user input was converted into text first then read out in a Benoit Blanc speech model. I'm thinking something like that only less amusing.