Earlier quoted context omitted.
There are really significant, novel copyright issues implicated by these large generative models trained on other people’s IP. If you take a step back, you can see that there are different ways to frame what is happening. One frame is: “Defendant built an algorithm that memorized features of Plaintiff’s IP. Defendant’s algorithm recombines parts of those features in order to produce works in the same domain that comp…
Artist today use the exact same method of learning from other peoples artwork to generate new artwork and styles. These models are learning just like any artist learns and then producing new content.
If you tried learning, let's say, the chiaroscuro technique from Caravaggio you'd be analyzing the way the painter simulated volumetric space by using white and dark tones in place of natural lighting and shadows. You wouldn't even think of splitting the whole painting into puzzle size pieces while checking how many how those look similar when put close one another.
Given somewhat decent painting skills, you'd be able to steadily apply this technique for the rest of your life just by looking at a very small sample of Caravaggio's corpus.
On the other hand if you tried removing even just a single work from the original Stable Diffusion data set you used to generate your painting, it would be absolutely impossible to recreate a similar enough picture even by starting from the same prompt and seed values.
Given how smart some of the people working on this are, I'm starting to believe they're intentionally playing dumb to make sure nobody is going ask them to prove this during a copyright infringement case.