Earlier quoted context omitted.
As an ML researcher, I agree. Meta doesn't include adequate information to replicate the models, and from the perspective of fundamental research, the interest that big tech companies have taken in this field has been a significant impediment to independent researchers, despite the fact that they are undeniably producing groundbreaking results in many respects, due to this fundamental lack of openness This should als…
Which option would be better? A) Release the data, and if it ends up causing a privacy scandal, at least you can actually call it open this time. B) Neuter the dataset, and the model All I ever see in these threads is a lot of whining and no viable alternative solutions (I’m fine with the idea of it being a hard problem, but when I see this attitude from “researchers” it makes me less optimistic about the future) > a…
We can't prove that a model like llama will never produce a segment of its training data set verbatim.
Any potential privacy scandal is already in motion.
My cynical assumption is that Meta knows that competitors like OpenAI have PR-bombs in their trained model and therefore would never opensource the weights.