My biggest worry is the idea of generating data as training data. We're obviously already unwittingly doing this, but once someone decides to augment low-volume segments of the dataset with generative input, we're going to start getting some really crappy feedback loops.
Do you ever have novel ideas while walking or in the shower? Well, you’re learning from data you generate.
These new text books could be great at simplifying the subject matter and making the material accessible or they may just never have fully understood the materials and are misleading.
Now imagine that over and over again, imo it's pretty likely to introduce inaccuracies if just taking a naiive approach.