> that just means the person who wrote such a summary needs to instead sue for copyright infringement right?
Doesn't need to be a person, could be another AI that wrote the summaries. I see a big problem for copyrights looming on the horizon - LLMs can reword, rewrite or generate input-output pairs using copyrighted data as reference, thus creating clean data for training. AI cleanly separates knowledge from expression. And maybe it should do so just to reduce inconsistencies and PII in organic text.
Copyrights should only be concerned with expression not knowledge, right? Protecting knowledge is the object of patents, and protecting names the object of trademarks. Copyright is only related to expression otherwise it would become too powerful. For example, instead of banning reproduction of this paragraph, it would also cover all its possible paraphrases. That would be like owning an idea, the "*" version, not a unique sequence of words.
Does it even make sense to talk about copyrights when everything can be remade in many ways so easily? Copyright was already suffering greatly since zero cost copying became a thing, now LLMs are dealing the second blow. It's just a fig leaf by now.
If we take a step back, it's all knowledge and language, self replicating memes under an evolutionary force. It's language evolution, or idea evolution. We are just supporting it by acting as language agents, but now LLMs got into the game, so ideas got a new vector of self replication. We want to own this process piece by piece but such a thing might be arrogant and go against the trend. Knowledge wants to be free, it wants to mix and match, travel and evolve. This process looks like biology, it has a will of its own.