Earlier quoted context omitted.
Outliers are the reason databases exist. Any "average" is simply readily apparent, therefore irrelevant for serious in depth analysis. Adding noise and fuzzing has a long history in statistics since the '70s [1], and while it does work on large numbers, it almost always messes up the details ie. the error bars. C.D. DP is essentially a cheap ripoff of the ideas implemented in ARGUS[2]. [1] 1977 Dalenius, see Do Not F…
> Outliers are the reason databases exist. Disagree. Data is why databases exist. > Any "average" is simply readily apparent, therefore irrelevant for serious in depth analysis. I said "aggregate", not "average". There are many kinds of aggregate analysis useful (in Astrophysics, you can take many different samples from different stars and use the aggregate to compute commonalities in the sample that you would've det…
But as you say: your "aggregate analysis" NEEDS "many different samples from different stars". Commonality is the result of your analysis based on different samples. But since they are common, you can go and sample and have the result without doing mass surveillance on every star.
ps: I am fully aware of photo stacking, but also note, that stars are not humans, see context of privacy. Please look at argus or sdcMicroGUI from CRAN to get a feeling for data utility vs. reidentification risk.