> However, they tend to memorize unwanted information, such as private or copyrighted content, I mean humans don't forget copyrighted information. We just typically adjust it enough (some of the time) to avoid getting a copyright strike while modifying it in some way useful. We don't forget 'private' information either. We might not tell other people that information, but it still influences our thoughts. The idea of…
I agree. As far as copyrighted and artistic works go, I've never fully understood what the objection is. If the work is being remixed not copied then it surely falls under fair use? Meanwhile, if it creates something new in an artist's style, it's only doing what talented imitators routinely do. There's the economic argument. But if that's accepted, then for fairness it would have to be extended to every other profes…
IMO the only reason there's even a question about whether LLMs can legally be trained on copyrighted works without permission is that the training is being done by (agents working on behalf of) rich people. If you or I scraped up every copyrighted work we could get our hands on without ever asking permission, trained an LLM on it, and then tried to sell access to the result? Just ask Aaron Swartz how that sort of thing goes, and his actions were orders of magnitude less.
Humans don't forget copyrighted material but we also don't normally memorize it. It takes substantial time and effort to be able to reproduce copyrighted material with just your brain.