Earlier quoted context omitted.
My guess is the big players hope is to steal an enough content and then build a self training LLM based off synthetic content (rehashed original works) before the theft part matters. Not sure how close they are to achieving but this seems to be a common SV gamble. Steal or do something shady, raise enough money / power so by the time your noticed, you have the money to win in the courts. You already see the propagand…
An LLM can be trained to find relevant knowledge online. It doesn’t have to be trained on the entirety of all existing knowledge.
Why do you think chatGPT lost its Web Search plugin lately? Copyright lawsuits. You can't even use copyrighted content in the prompt because it will make the model makers liable.