>
the popular public ones are mostly trained on stolen/pirated texts offthe internetYou mean like actual literature, textbooks and scientific papers? You can't get them in bulk without pirating. Thank intellectual property laws.
> from social media clouds the companies control
I.e. conversations of real people about matters of real life.
But if it satisfies your elitist, ivory-towerish vision of "healthy information diet" for LLMs, then consider that e.g. Twitter is where, until now, you'd get most updates from the best minds in several scientific fields. Or that besides r/All, the Reddit dataset also contains r/AskHistorians and other subreddits where actual experts answer questions and give first-hand accounts of things.
The actually important bit though, is that LLM training manages to extract value from both the "bullshit" and whatever you'd call "not bullshit", as the model has to learn to work with natural language just as much as it has to learn hard facts or scientific theories.