Earlier quoted context omitted.
You took that quote out of context and missed the broader point in the process. The snippets provided in regular search results cannot generally replace the substance of the full articles they link to, while that's the whole point of GP's hypothetical website—it simply doesn't reproduce large chunks of text verbatim, presumably to avoid copyright infringement claims in the hypothetical's frame, and in GP's rhetorical…
>content whose substance was created by someone else And how did the training data contribute to the content in any meaningful way? Inspiration isn't substance. You think all fantasy writers gotta pay Tolkien estate bc so much of fantasy draws from his tropes? Lmao no.
If training data is so unimportant, why not simply not use it and avoid the controversy? At the very least that would certainly fix the issue where the model demonstrates how "inspired" it is by NYT articles by reproducing them verbatim.
:)