Earlier quoted context omitted.
All of them are trained on copyrighted data. Why is it okay for a model to serve up paraphrased books but verbatim copies from the Pirate Bay are illegal? I don’t deny the utility of LLMs. But copyright law was meant to protect authors from this kind of exploitation. Imagine instead of “magical AGI knowledge compression”, instead these LLM providers just did a search over their “borrowed” corpus and then performed a…
> Why is it okay for a model to serve up paraphrased books but verbatim copies from the Pirate Bay are illegal? Because they are not actually memorizing those books (besides few isolated pathological cases due to imperfect training data deduplication), and whatever they spit out is in no way a replacement for the original? Here's some back-of-the-envelope math: Harry Potter and the Philosopher's Stone is around ~460K…
Is this actually true? I think in many cases it is a replacement. Maybe not in the case of a famous fictional work like Harry Potter, but what about non-fiction books or "pulp" fiction?
Kind of feels like the bottom rungs of the ladder are being taken out. You either become J.K. Rowling or you starve, there's no room for modest success with AI on the table.