> People are going to fight hard against this, partly because it’s annoying and partly because of (imho exaggerated) patient privacy related concerns. Somebody’s going to try make some kind of gated thing where you have to prove you have a PhD and a “legitimate cause” before you can access the data, and that person should be fought tooth and nail...
I have a huge amount of respect for Scott, and the rational thinking that he shares with us. However, I'd like to make two arguments against the above.
1. Personally-identifiable information is protected by law over on this side of the pond. This is one of those cases where the thinking is quite different between the US and Europe. In Europe, you need to get the permission of everyone involved in the study before you can publish your data about them, if there is any chance that they can be identified from that data. In a lot of ways this is good, but one way in particular is that it encourages pre-planned proper studies, where you set it all up properly and get every participant to sign the permission sheet, which is good for science.
2. Sometimes absolutely full data release is overkill and inappropriate. For example, if I'm planning to publish a paper saying that mutations in a particular small region of the human genome cause a particular disease, then it is sufficient to publish the list of causative mutations. Full data release would mean publishing the entire sequencing run for the patient - that's inappropriate because you can identify someone from it, and you also then know a huge amount of information about them - their traits and congenital conditions. A lot of journals are pushing for this, and the compromise that we seem to have reached is to store the full sequencing data in a secure repository and provide access given a really good reason. Likewise, a census will report statistics about areas of a country, but the raw data is kept back because it is deeply personal in nature.
So my points are that data shouldn't go into a study unless it can be openly published, and that there are types of raw data that are too personal or voluminous to openly release, and the data derived from it should be released instead.