Earlier quoted context omitted.
[flagged]
I am not trolling. I don’t recommend engaging in a good-faith dialogue with someone by prefacing it with your belief that they are trolling. If you're basing your trust in ChatGPT on the claim that it is trained on Wikipedia, you might as well read Wikipedia instead because then you also see the sources for each claim, or the fact that certain claims are unsourced. ChatGPT will not let you know if a certain claim is…
Digression aside, I would impress upon you I haven't claimed, nor do I believe what you're claiming I have. I am simply providing you with a slice of the larger original academic paper that provides great, rigorous, and peer-reviewed documentation of precisely what GPT is, and what it's trained on.
From this paper you'll perhaps gather an instinct that these folks working on this problem are widely aware of the extremely well-known concept of "open source knowledge" and were well considered in their application of pruning data to their needs.
I believe you'll perhaps further gather retrospective insight upon the idea that GPT is doing anything more than giving you a T9-predictive-autotext for the entirety of that dataset; meaning you can try to coax it into saying anything you want but if the p-values aren't right or the predictive potential of a given token "coming up next" isn't there, that's just.. how it goes.
It's data. It's not quite the tower of Babel, but we'll get there soon enough. Kind of like how all the fancy 3D video game rendering software that looks incredible these days is still just manipulating tuples and vectors with matrix calculations, stuff you could do on paper but why would you do that math to describe a picture when you could just draw it.
ChatGPT is just drawing the pictures (in this poorly chosen analogy), and giving us the cool graphics. The math is all pretty benign and based in the fundamentals of neural networks, not even the more fanciful CV and deeper trained NLP can get to.
Info-scientists, I wonder what the "rainbow table" of all language and ideas etc. that would be relevant to a latent language learning model would be... this is a wonder to me because I lack the sufficient knowledge and field expertise.