Earlier quoted context omitted.
Not at all, I love talking about it. I was convinced that a model needed to be able to read like we do. And what do we do when we read? Pick up a book. That turns out to be surprisingly hard, at least for training data. Step one is to acquire the books. Step two is to turn them into a readable format for computers. Both steps were very hard. I lucked out on step one because The Eye happened to host all of bibliotok,…
oh wow, that was a surprisingly awesome story. thanks for sharing! this maaaay be covered in the Pile's writeup (which i have not yet read) but i wonder who was curating the overall "mix" of the content. seems easily biased to, say, public domain books, since the corpus is easily available. when people say things like "GPT3 has been trained on all of the internet" i suspect this is a gross exaggeration. In reality it…
Also bmk. https://twitter.com/nabla_theta?s=21&t=Gt6YrATJHnmY046MdzhYD...
They did the legwork of writing the paper and getting everything into a presentable format. A bunch of other people helped too; I wasn’t as involved as I could’ve been.
It was all discord-based. As far as I know it was the first serious research collaboration to happen solely via chatroom.
bmk also got the 50GB of code from GitHub, I think. So that’s where GPT-J’s coding ability likely came from.