I personally do not want the companies to release training data (at least for a while) because then it gives people leverage to neuter it. I don't want a sanitized LLM, and I don't have $60M lying around to train my own. Copyrighted material, sexual content, political opinions, throw it all in and release it please! Yes, reducing bias in the models is a noble goal, but introducing new bias and blindspots to do it is…
It's also a necessary goal in order for these models to be more broadly adopted.
We've seen too many examples of bias in the training data set manifesting in ways that actively discriminate against people. Which is unethical and in many places illegal.
And having copyrighted material and sexual content in your model will simply open you up to lawsuits as is happening right now between authors and OpenAI. Not sure that is a position most startups want to be in.