Interesting. When it's copyrighted works in a digital form (plain text, ePub, whatever) it's a legal issue. ROT13 the text is the data still a copyright violation? I imagine so since it's a trivial thing to restore it to a legally volatile form. I understand that an unresolved issue is whether, once ingested into an LLM, the trained LLM is in violation of copyright. One wonders if human readers too are in violation o…
At least it's a thing for vision models and was also used to train OpenAI Five (the dota AI)[0]. So it probably applies to LLMs too.
[0] https://cdn.openai.com/dota-2.pdf (page 2)