I'm a programmer in the GP data analysis world. We use the term 'pseudonymization' for this kind of data. 'Anonymization' is used solely to refer to, say, 'the sum total of diabetes patients this practice has' (that would be anonymous patient data; it would not be anonymous relative to the GP office this refers to): Aggregated data that can no longer be reduced to a single individual at all. The term raises questions…
Algorithm can pick out almost any American in supposedly anonymized databases
21–30 of 101 posts
Re: Algorithm can pick out almost any American in supposedly anonymized databases
#22Would differential privacy fix this problem? I heard that new US census will use it.
There are caveats. The exact strength of the privacy guarantee depends on the parameters you use and the number of computations you do, so simply saying "we use a differentially private algorithm" doesn't guarantee privacy in isolation.
Re: Algorithm can pick out almost any American in supposedly anonymized databases
#23In any database whose fields have at least 4 meaningful ranges, 16 fields give 4.000.000.000 possibilities. Now I am on my iphone. As long as ranges are meaningful (i.e. they do divide the group in somewhat even parts), the individuating possibilities are HUGE. And fields do have usually many more than 4 ranges. Anonymizing datasets is a weasel term. The database is secure or it is not. As any database is quite likel…
Re: Algorithm can pick out almost any American in supposedly anonymized databases
#24I'm a programmer in the GP data analysis world. We use the term 'pseudonymization' for this kind of data. 'Anonymization' is used solely to refer to, say, 'the sum total of diabetes patients this practice has' (that would be anonymous patient data; it would not be anonymous relative to the GP office this refers to): Aggregated data that can no longer be reduced to a single individual at all. The term raises questions…
It doesn't roll off the tongue, perhaps pseudo-anonymization is enough.
Pseudonimization is bad terminology in that it's indistinct from the above, to the point that parent has already mixed the two up while in the process of recommending it. And it'd be worse verbally.
"Pseudo-anonymization" could work, but something like "breakable anonymization" or "partial anonymization" might be better in that it's more obvious to a reader and doesn't rely on familiarity with technical terminology to convey the idea.
I'd go with breakable, myself, since it's most to the point about why it's a problem.
Pseudo is etymologically correct, but that doesn't necessarily help us much when the goal is ratio and ease of understanding by a wide population of readers.
Partial could work in the sense that you did part of the job, which people would hopefully understand is a bit like having locked the back door for the night while leaving the front propped wide open.
And there are probably other good options. If I was writing about this topic often, I'd strongly consider brainstorming a few more and running a user test where I ask random people to explain each term, then go with what consistently gets results closest to what I'm trying to discuss.
Re: Algorithm can pick out almost any American in supposedly anonymized databases
#25I'm a programmer in the GP data analysis world. We use the term 'pseudonymization' for this kind of data. 'Anonymization' is used solely to refer to, say, 'the sum total of diabetes patients this practice has' (that would be anonymous patient data; it would not be anonymous relative to the GP office this refers to): Aggregated data that can no longer be reduced to a single individual at all. The term raises questions…
Pseudo and psuedo are derived from the greek work ψεμα psema, which literally means lie. It's synonym is 'untrue' which indicates that your term is the right term to use when anonymizing-but-not-really-anonymizing
Re: Algorithm can pick out almost any American in supposedly anonymized databases
#26I'm a programmer in the GP data analysis world. We use the term 'pseudonymization' for this kind of data. 'Anonymization' is used solely to refer to, say, 'the sum total of diabetes patients this practice has' (that would be anonymous patient data; it would not be anonymous relative to the GP office this refers to): Aggregated data that can no longer be reduced to a single individual at all. The term raises questions…
Re: Algorithm can pick out almost any American in supposedly anonymized databases
#27Would differential privacy fix this problem? I heard that new US census will use it.
Yes, in the sense that the output of a differentially private protocol has mathematical guarantees against re-identification, regardless of the computational power or side information an adversary has. There are caveats. The exact strength of the privacy guarantee depends on the parameters you use and the number of computations you do, so simply saying "we use a differentially private algorithm" doesn't guarantee pri…
Re: Algorithm can pick out almost any American in supposedly anonymized databases
#28Earlier quoted context omitted.
It doesn't roll off the tongue, perhaps pseudo-anonymization is enough.
Pseudonymization already refers to reference by pseudonym. Pseudonimization is bad terminology in that it's indistinct from the above, to the point that parent has already mixed the two up while in the process of recommending it. And it'd be worse verbally. "Pseudo-anonymization" could work, but something like "breakable anonymization" or "partial anonymization" might be better in that it's more obvious to a reader a…
Re: Algorithm can pick out almost any American in supposedly anonymized databases
#29The great contribution of Differential Privacy theory is to quantify just how little use you can get out of aggregate data before individuals become identifiable. Unfortunately, Differential Privacy proofs can be used to justify applications which turn out to leak privacy when the proofs are shown to be incorrect after the fact, when the data is already out there and the damage already done. Nevertheless, it is instr…
You are right that some differential privacy proofs have later been found to be wrong. For example, there is an entire paper about bugs in initial versions of the sparse vector technique [1].
However, I imagine this will evolve the way cryptographic security has evolved: at some point, enough experts have examined algorithm X to be confident about its differential privacy proof; then some experts implement it carefully; and the rest of us use their work because "rolling [our] own" is too tricky.
Re: Algorithm can pick out almost any American in supposedly anonymized databases
#30I'm giving less and less of a fuck about this shit. Mostly because there's a point at which you refuse to be put on the run and don't want to hide. So maybe this is where the fight begins. At some point people make up their mind about dying not being a very big deal, and to become masters of their fate. To die on your feet, looking your accusers in the eye. Who would be so bold as to accuse me? Could I dare entertain…