Earlier quoted context omitted.
Computers aren't humans and LLMs aren't human brains. We have no way to reconstruct memories from a preserved brain (yet). The exact ways in which humans form memories and store information isn't even known yet; we're still drilling into the specifics from higher-level concepts. Modeling the human brain like nodes with weights ignores a lot of biological processes. Blood/oxygen flow, hormones, neurotransmitter decay,…
So if we add a bunch of complex processes to an LLM, in order to produce a better analog of a human in terms of degree of complexity if not actual function, does that have some bearing on this copyright question? It doesn’t seem clear to me that it does. Is the argument that sufficient complexity in how an “intelligence” processes this copyrighted data leads to the output being transformative vs not transformative in…
If you turn an LLM into a person, you may have a ethical and legal basis for treating that LLM like a person. There's no law about artificial intelligence being sentient or not, but law applies only to people, so that'd be the supreme court case of the century. I remember the Star Trek TNG episode about this topic and while the answer was perhaps more obvious with mister Data, the best arguments for and against synthetic consciousnesses have all been made in that episode.
The law doesn't care for how human-like programmers may think their program is, and neither should it in my opinion. What matters to the law is that the output is a result of an automated process, which comes with a completely separate set of rules and conditions compared to fair use.
The complexity of the program isn't a very good legal defence in my opinion because there's no clear line when the complexity would be enough to be considered human like. You could, for example, also claim that a computer is just very good at doing imitations, just like a person can be good at doing imitations on stage, and that an mp3 file is just an elaborate imitation act.
Even with a digital system identical to a physical system I don't think you can state human-ness as an argument. A tape recorder is just a sophisticated, automated way of sending an electric field through a magnet, similar to what a human can do with a dynamo and a spool of tape; a sort of delayed-action theremin, which would turn it into a musical signal. In turn, a neural network can be solved by human brain power if you pay enough people to work on a single iteration for an entire year. Almost everything a computer can do is just a sophisticated way of doing what humans are already doing, so I don't see why this is different when it comes to this topic.
There's no obvious "this is human like now" threshold and I doubt there will be until we know exactly how the human brain works.
When it comes to AI generated works, we don't currently know where the line between copyright violation and copyright exemption lies. If this goes through, it's the third major lawsuit of its type, the other two being actions against Stable Diffusion by artists to prevent it from producing derived works from their art.
IANA but I think you can assume that the "but computers are just like digital humans" approach won't fly in court. That's not really important, though; both sides of the coin are already having clever lawyers write up legal defences for their points of view, and it's more than likely that some other factors ("is a model a derived work" (probably) and "is the output of a model a derived work" (who knows!)) will decide the future of AI and copyright. The interesting thing is that academic research is essentially exempt from copyright law, so nobody can demand takedown of their content from research data sets, but whether the commercial branch of AI companies can use their academically generated models to serve their customers?
As an upside, I don't think the death of ChatGPT is the end of AI. This whole scenario could've easily been avoided if AI companies paid for their data set or had restricted themselves to works they had the license to (public domain, CC0, etc.), and OpenAI in particular has been pretty brazen in their "we'll see about it if it ever comes up" approach. Companies like Github are probably in the best place legally, where their users are already signing off on "we can take your content and do whatever the fuck we want" terms and conditions.