Earlier quoted context omitted.
Wouldn't that depend on the use case? If you just had the model regenerate articles that roughly approximate its source material that is much a more clear cut violation of a paywall. But if you use that data as general background knowledge to synthesize aggregative works such a history of the vietnam war, or trends in musical theatre in the 1980s relative the 1970s, or shifts in the language usage of formal honorific…
Yes, I think this is a rather fact-specific inquiry. My main point is that the research/commercial distinction is not the only factor (and not even the most important one). > if you use that data as general background knowledge to synthesize aggregative works such a history of the vietnam war, or trends in musical theatre in the 1980s relative the 1970s, or shifts in the language usage of formal honorifics, then that…
I think the question is if it changes analysis if the dataset DOES include a bunch of books and articles related to Vietnam beyond your specific book.
In the first cast where it just rewriting a single books content, the unfairness is clear.
But in case where it is producing a new synthesis and analysis of the data, derived in part, but not regurgitating the source material, is that unfair?
The latter isn't clear to me.