Earlier quoted context omitted.
It would be incredible for LLMs. Searching it, using it as training data, etc. Would probably have to be done in Russia or some other country that doesn't respect international copyright though.
Do you have a reason to believe this ain't already being done? I would assume that the big guys like openai are already training on basically all text in existence.
https://torrentfreak.com/meta-torrented-over-81-tb-of-data-t...