That's pretty cool. When I tried stuff like that, it was entirely based off of the prompt I gave it. I asked it to criticize my HN posts, and it came up with a story about how I was a charlatan; I asked it to summarize my views on Python exceptions in my HN posts, and it came up with a story about how I'm a world renowned Python expert. (I'm neither, but have written nothing noteworthy, so it's unsurprising.)
It could maybe be added to the detector but I think we're going to discover "did this output come from this model?" is ultimately undecidable. My intuition is that this is somehow equivalent to the halting problem, but that's a wild guess, and I'm not an AI researcher or computer scientist (just a humble SWE). Something like, if you start generating prompts that might take you to a part of latent space where this output could have been produced, you don't know when to stop generating longer and longer prompts.