Live data from Hacker News

How Cambridge Analytica’s Facebook targeting model really worked

niemanlab.org

71–80 of 210 posts

Re: How Cambridge Analytica’s Facebook targeting model really worked

#71
post #68
post #41

Spoiler warning. Article punchline ahead. "The whole point of a dimension reduction model is to mathematically represent the data in simpler form. It’s as if Cambridge Analytica took a very high-resolution photograph, resized it to be smaller, and then deleted the original. The photo still exists — and as long as Cambridge Analytica’s models exist, the data effectively does too." That's an eloquent piece of explanati…

> when strictly speaking the raw data has indeed been deleted after being used to create a derivative work that can for all important purposes be used to recreate the original? To be precise, you almost certainly cannot use this data to recreate anything remotely resembling the original dataset. This type of dimensionality reduction would throw away enormous volumes of data. There is no meaningful sense in which you…

> It's arguable whether they should be allowed to keep those insights, but there's no privacy risk there really.

So if Google has distilled someone's emails over the years into "closeted homosexual with a deeply repressed leather fetish", that's not an invasion of their privacy as long as they throw away the source materials?

Re: How Cambridge Analytica’s Facebook targeting model really worked

#72
post #68
post #41

Spoiler warning. Article punchline ahead. "The whole point of a dimension reduction model is to mathematically represent the data in simpler form. It’s as if Cambridge Analytica took a very high-resolution photograph, resized it to be smaller, and then deleted the original. The photo still exists — and as long as Cambridge Analytica’s models exist, the data effectively does too." That's an eloquent piece of explanati…

> when strictly speaking the raw data has indeed been deleted after being used to create a derivative work that can for all important purposes be used to recreate the original? To be precise, you almost certainly cannot use this data to recreate anything remotely resembling the original dataset. This type of dimensionality reduction would throw away enormous volumes of data. There is no meaningful sense in which you…

Analogy does not work here and is misleading. You cannot do much if anything with 20 most representative pixels (if there is such a thing) but you can infer highly valuable characteristics about the person. Yes, you cannot recreate the original data but what you end up is potentially much worse (sensitive/private) than the original data.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#73

Earlier quoted context omitted.

> The people seeing these ads are under the assumption that everyone else sees them, not that it's specifically targeted at their personality type. How long will that be true? Do people make that assumption about search results?

Outside of the tech bubble, simply saying "yes" would be disingenuous. They're not even asking the question in the first place

I think retargeting has thoroughly blown up the idea that ads online are shown to everyone. My non-technical acquaintances are very aware of why certain products follow them around the internet in ads.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#74
post #5

Earlier quoted context omitted.

It's a hard problem but Netflix's model represents the state of the art in machine learning for recommendations. Still scared of the singularity? :)

Really? Because in terms of actual usefulness, I find youtube's suggestions to better...

I tend to agree with you. The only thing I dislike about it is when I happen to open a link I'll randomly be sent or notice in an article that's something unlike what I normally enjoy or something I close right away and then for the next few days, or until I watch a bunch more of what I usually enjoy, the entire suggestion list is only things to do with that one random link.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#75
post #71
post #68

Earlier quoted context omitted.

> when strictly speaking the raw data has indeed been deleted after being used to create a derivative work that can for all important purposes be used to recreate the original? To be precise, you almost certainly cannot use this data to recreate anything remotely resembling the original dataset. This type of dimensionality reduction would throw away enormous volumes of data. There is no meaningful sense in which you…

> It's arguable whether they should be allowed to keep those insights, but there's no privacy risk there really. So if Google has distilled someone's emails over the years into "closeted homosexual with a deeply repressed leather fetish", that's not an invasion of their privacy as long as they throw away the source materials?

If they kept information like that, then yes that would be an invasion of privacy. But that sort of information is almost certainly not encoded in an ML model trained on 50 million people's data.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#76
post #72
post #68

Earlier quoted context omitted.

> when strictly speaking the raw data has indeed been deleted after being used to create a derivative work that can for all important purposes be used to recreate the original? To be precise, you almost certainly cannot use this data to recreate anything remotely resembling the original dataset. This type of dimensionality reduction would throw away enormous volumes of data. There is no meaningful sense in which you…

Analogy does not work here and is misleading. You cannot do much if anything with 20 most representative pixels (if there is such a thing) but you can infer highly valuable characteristics about the person. Yes, you cannot recreate the original data but what you end up is potentially much worse (sensitive/private) than the original data.

That's not really true, and is kind of a fundamental misunderstanding of how these things work.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#77

I am really puzzled by the Cambridge Analytica scandal. It's not particularly savory, but is there something happening here that it wasn't basically already known about how Facebook worked? By the protests of their own executive, the system was working as designed, and at worst Cambridge Analytica misled them about how they intended to use the data, right? There was no actual security breach here, as far as I can und…

It became a "problem" because it helped Trump win.

This is exactly it. At least it stops the news from droning on and on about Russia.

I thought Clinton spent large amounts of money on data and the Democrats admitted the data was bad or at least that was their excuse. How much did CA pay for this data? I still find it crazy that Trump campaign spent 30% of what Hillary did and still won. The Russians used 100k$ worth of ads to sway the election. This stuff doesn't t add up.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#79
post #71
post #68

Earlier quoted context omitted.

> when strictly speaking the raw data has indeed been deleted after being used to create a derivative work that can for all important purposes be used to recreate the original? To be precise, you almost certainly cannot use this data to recreate anything remotely resembling the original dataset. This type of dimensionality reduction would throw away enormous volumes of data. There is no meaningful sense in which you…

> It's arguable whether they should be allowed to keep those insights, but there's no privacy risk there really. So if Google has distilled someone's emails over the years into "closeted homosexual with a deeply repressed leather fetish", that's not an invasion of their privacy as long as they throw away the source materials?

That's a ridiculous response. If they managed to infer this characteristic from emails, what they would keep is a tool which, given that set of emails again, infer the same characteristics (and theoretically a similar set of emails). They would by no means be allowed to keep the kind of information you described.

What is more relevant is a model which, given characteristics such as "closeted homosexual with a deeply repressed leather fetish", they would be able to infer other characteristics, such as support of particular political candidates, responsiveness towards targeted political or commercial ad campaigns, etc. That's what's relevant here.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#80
post #75
post #71

Earlier quoted context omitted.

> It's arguable whether they should be allowed to keep those insights, but there's no privacy risk there really. So if Google has distilled someone's emails over the years into "closeted homosexual with a deeply repressed leather fetish", that's not an invasion of their privacy as long as they throw away the source materials?

If they kept information like that, then yes that would be an invasion of privacy. But that sort of information is almost certainly not encoded in an ML model trained on 50 million people's data.

How do you know what CA trained on, or what's possible? Do you have qualifications in ML?
Post reply on HN