I'm not sure how this is an attack. Is it actually vital that models don't repeat their training data verbatim? Often that's exactly the answer the user will want. We are all used to a similar "model" of the internet that does that: search engines. And it's expected and required that they work this way. OpenAI argue that they can use copyrighted content so repeating that isn't going to change anything. The only issue…
Has anyone done any work to produce citations for the generated data?