I personally do not want the companies to release training data (at least for a while) because then it gives people leverage to neuter it. I don't want a sanitized LLM, and I don't have $60M lying around to train my own. Copyrighted material, sexual content, political opinions, throw it all in and release it please! Yes, reducing bias in the models is a noble goal, but introducing new bias and blindspots to do it is…
We really need to somehow separate a bias towards accuracy as distinct from some bias towards say, a sports team. Using everything would be like taking a bunch of students final exams and then claiming the most common answers are the correct ones. This isn't how expertise and accuracy works. Most things worth doing are not only genuinely hard and complicated but something that only a minority subset of accomplished p…
Humans will argue the right answer until their last days. It's frustrating how on-the-fence chatgpt can be. It's pretty interesting too, because in a professional environment one of the most important things you need to do is have an opinion and take a position otherwise you can't execute.