About the training data, cant the datasets from the Tulu3 Model by the Allen Institute be used? They claim that they have used a fully open source training dataset.
If a collective/coop of individuals and organizations with storage and network capacity could collaborate with each other to archive and index deduplicated training data that would be huge.
Perhaps this is already happening. I was looking at Red Pajama last year as an example.
Someone like myself could arrange to host 200+TB on high speed storage with a 10G public IP for example, then we get a bunch of us together and hopefully access to training datasets would be decentralized and uncensored in an idea setup.
Is all that in progress and I just need to learn how to join?
Is Red Pajama something to look at again?
Is there someone tracking datasets in detail like HuggingFace has all the models? I know a lot of datasets are on it also, but there is massive duplication.