Playing with the AI dungeon a while back (on the GPT-2 mode) I was presented with a tilapia recipe, titled "Kittencal's Broiled Tilapia" - it sounded bizarre so I decided to do a google search and I found that it was directly pulled from from https://www.recipezazz.com/recipe/broiled-parmesan-tilapia-7... - the user who posted it was 'Kittencal'
So in terms of copyright, is GPT-2 a derived work of that recipe? Or generally are models derived works of their training data? It seems lots of people use training data from Flickr, like COCO, and then use the resulting model for commercial services.
That said, now a lot of privacy policies make sense.
ALSO, it makes me wonder about "Your call is recorded for training purposes"... is it a coincidence or is it very carefully worded?