Earlier quoted context omitted.
> do you think they are just lying Yes. They're constantly lying, and constantly getting caught for it. They have a reputation for it. Why do you think this would be any different? Their standard opt-out agreement frames it as if they won't train on your data, but they do anyway, due to legal loopholes. They essentially clean-room everyone who opts-out, so while it's "technically" not training on "your" data, to the…
Can you elaborate on the loopholes here?
Papers like "Curated Synthetic Data Doesn't Have to Collapse" [1] and "How to Synthesize Text Data without Model Collapse?" [2] demonstrate it's possible to do this.
Since OpenAI's Privacy Policy [3] explicitly allows for the use of deidentified data, it's possible they consider rewrites (maybe paired with a model used to identify explicit PII) to be deidentified. Whether OpenAI's legal team thinks rewriting in this way technically means they aren't training on your data isn't something I'm able to comment on.
Here's the relevant Privacy Policy statement:
We also aggregate or de-identify Personal Data so that it no longer identifies you and use this information for the purposes described above, such as to analyze the way our Services are being used, to improve and add features to them, and to conduct research. We will maintain and use de-identified information in de-identified form and not attempt to reidentify the information, unless required by law.
Please note all the hedging words I used (maybe, possibly, etc). I honestly have no clue if they are doing this. I'm merely elaborating on a possible loophole like you asked.[1]: https://arxiv.org/abs/2605.07724