Ah sorry, I was trying to say two things at once referring to two different contexts. In both cases the only 'temporary copying' takes place.
- In the context of specifically training an AI-model under EU law: -
Article 5(1) of the Directive 2001/29/EC [1] is argued to apply. Of course, if sued, AI companies still need to show that their use is otherwise lawful (which should be fairly easy, since they actually do not retain anything remotely resembling a copy at all)
- In the context of specifically training an AI-model under US law: -
The purpose of the copy can be argued to be non-exploitative, incidental, for the purpose of enabling technology, and temporary.
By contrast: Google Books even argued that permanently retaining entire copies of books wholesale was fair use [2], provided they didn't provide copies of those books to 3rd parties.
OpenAI argues that there's no way they're doing anything even remotely close to that. [3]
Disclaimer: IANAL, YMMV.
[1] https://eur-lex.europa.eu/LexUriServ/LexUriServ.do?uri=CELEX...
[2] https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....
[3] https://www.uspto.gov/sites/default/files/documents/OpenAI_R...
See also:
[4] https://news.ycombinator.com/item?id=37879938 Previous thread where I looked up a bunch of sources.