I was hoping the article would propose the opposite: if you train LLMs on copyrighted data, you owe the author a part of your income from it. How big should be determined by courts but probably proportional to the amount of data. There's absolutely no reason rich people owning ML companies should be getting richer by stealing ordinary people's work. But practicality trumps morality. The west needs to beat China and C…
Realistically, the copyrighted works that are most valuable to training machine learning models, at least if we go by The Pile as typical of training data [1] is: - Web pages; hard to argue that royalties are due since these are publicly available for free - Scientific papers; these do cost money but the copyright is typically owned by scientific publishers - Github, Stack Exchange, HN (yes); these are freely availab…
If you're going to ignore the existence of copyright and licenses we should extend it to everything that's ever been posted on the internet, not just "web pages". Why shouldn't all books and films count as free too?
I'm actually open to the idea of just abolishing copyright but it's kind of silly to act like it's only about Elsevier. Lots of creatives depend on copyright in order to earn a living, similarly to how patents fund a lot of important research despite how noxious the patent system has become.
If we fixate on examples like Elsevier or Martin Shkreli in order to argue for completely abolishing the copyright or patent systems we risk destroying the framework that enables valuable creative works or new technologies to be developed in the first place. This is part of why people are so upset by AI companies arguing that they should just be able to ignore the whole framework in order to enrich themselves; once you allow the for-profit AI companies to do it, other groups are going to line up to also demand a free ride.