It's so obvious to me that machine learning models are derivative works of their training set. If they weren't, then why would these companies fight so hard to say otherwise? They need that training data to make their product, so they should pay the licensing fees for it! 10 years ago, when I worked on a machine learning model for my employer, it was unthinkable to train on data we did not have the rights to use. But…
> It's so obvious to me that machine learning models are derivative works of their training set. Okay, but narrative creators watch movies and listen to music and read books too. Many do indeed "file the serial numbers off" other people's work and publish something else, that makes them money and not the original creators. Does one instance of "filing the serial numbers off" by one author mean that no authors anywher…
It is. It obviously is. It's the same reason that a person watching a movie and remembering it later is different than recording the movie with a camcorder.
> Ah but I made a robot that walks into theaters, buys a ticket, records the movie, leaves, and then recreates the movie at my home an infinite number of times. I didn't break the law, since a human could surely do the same thing with enough practice and effort.
Do you realize how ridiculous that sounds?