Is there any details on the input training datasets? Slight off topic but if anyone knows.
Also curious... I spent a while trying to get system to tell me directly, but no dice: https://twitter.com/nickmvincent/status/1598478685019189248?... It gives a generic answer that it's some proprietary combinations of "books, articles and websites". I'd guess Wikipedia is in there for sure (English and maybe other editions as well), something like "BookCorpus" ( https://huggingface.co/datasets/bookcorpus ), probabl…
https://en.wikipedia.org/wiki/Common_Crawl
I also have a an odd hunch ChatGPT might have used a scihub mirror as inputs for example.