Interestingly, doesn't include the text in TFA (doesn't cover copyright at all). Not sure if hallucinated or referring to a system prompt at another layer.
My understanding is that it's not actually XML, but special tokens that can be injected into the input stream (and read from output stream). The XML format is just how those tokens are sometimes rendered for human consumption.
"My understanding is that these APIs are not actually XML, but actually a bytestream. The XML format is just how it is usually rendered for human consumption."
At best this is a small piece IMO, but this does look like a piece of it, well done. At least based on my own experience exposing sys prompts from niche llm models in my saas niche. Should be higher upvoted. For those asking for proof, unfortunately it simply isn’t possible short of anthropic employee disclosure. All I know is that this is exactly how LLMs tend to leak: provide info they didn’t realize was sys prompt related and thus leak what they aren’t supposed to leak within their own current context window.