Earlier quoted context omitted.
They do: https://xcancel.com/vxunderground/status/1888019174133276846 , https://www.theverge.com/2023/7/9/23788741/sarah-silverman-o... The tweet only names Meta, but it would be very surprising if OpenAI didn't do the same thing.
Anyone who doesn't train on all material available, legal or otherwise, will be outcompeted by teams that do, including those based in countries that don't respect Western copyright law. It's that simple. Either this is practice is judged (or legislated) to be fair use, or copyright is done. It's also that simple.
Copyright law exists for a reason. Trying to improve an LLM doesn't give you the right to flout our legal system. Yes, other countries might have an advantage in LLM training as a result but so be it.