Live data from Hacker News

How Cambridge Analytica’s Facebook targeting model really worked

niemanlab.org

151–160 of 210 posts

Re: How Cambridge Analytica’s Facebook targeting model really worked

#151
post #89
post #76

Earlier quoted context omitted.

That's not really true, and is kind of a fundamental misunderstanding of how these things work.

Unless the data is completely random it's not crazy to say that the data can be reconstructed from a reduced version. If you have a million points that largely fall on a 3-dimensional line and you project that into 2 dimensions, you can easily recover that lost dimension with losses relative to the deviation. And that loss may not even matter depending on the kinds of data and margins of error you're working in.

This is actually a nice illustration of the central problem with this argument: the more personally identifiable a piece of information is, the less recoverable it'll be, and vice-versa. If all of the points of data are on some n-dimensional line, then obviously all of them can easily be recovered, but knowing all those things about a person doesn't actually tell you any more about them than knowing just one of those things. Conversely, if the points of data are very random then it'll only require a handful of points to uniquely identify a person and find the entry in the original data set with all their other information, but dimensionality reduction will have to throw that data away - you simply won't be able to recover that information from the model. (We actually know from the literature on de-anonymization that a lot of data falls into the second category.)

Re: How Cambridge Analytica’s Facebook targeting model really worked

#152
post #92

Very interesting article but I wish it went one step further. Why does it matter that cambridge analytica knew a user's big five or that they were an old, uneducated republican? How was this (inferred) data used? I assume they wrote/created different ads for different sets of users... but how many segments did they have? Did their graphic designer build 500 different ads, or was text/images dynamically inserted based…

I wonder if you learned something since your blatant failure last week where you gave all that trust into facebooks hands, and praised advertisements, one day before they've blown up.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#153
post #68
post #41

Spoiler warning. Article punchline ahead. "The whole point of a dimension reduction model is to mathematically represent the data in simpler form. It’s as if Cambridge Analytica took a very high-resolution photograph, resized it to be smaller, and then deleted the original. The photo still exists — and as long as Cambridge Analytica’s models exist, the data effectively does too." That's an eloquent piece of explanati…

> when strictly speaking the raw data has indeed been deleted after being used to create a derivative work that can for all important purposes be used to recreate the original? To be precise, you almost certainly cannot use this data to recreate anything remotely resembling the original dataset. This type of dimensionality reduction would throw away enormous volumes of data. There is no meaningful sense in which you…

> To be precise, you almost certainly cannot use this data to recreate anything remotely resembling the original dataset. This type of dimensionality reduction would throw away enormous volumes of data. There is no meaningful sense in which you can reconstruct the data from it.

First off, I think that's wrong. The idea is after all to keep the information that will result in the smallest error compared to the original on the dimensions one cares about. Within what the model emphasizes a reconstruction can be not only "remotely resembling the original dataset" but as closely resembling the original dataset as is possible with the capacity of the representation.

Next, I'm really not talking only about the particular method described in the post. It's definitely possible to choose to make a light enough reduction to preserve the aspects of the information one is interested in, and to optimize for recall rather than generalization. A more realistic context is going to be that some information about the affected individuals is still exposed or kept (maybe in a compact derived form), which would in many cases give excellent possibilities to restore information accurately enough that claims to have the removed the data are effectively deceptive.

Even for cases where the models are in good faith created only to "distill some insights" I'm skeptical that they really are useless for recovering individual information. I'm by no means an expert in differential privacy but I do listen when it comes up, and a lot of what we see from that field seems to come down to being able to trade off the relation between keeping the data useful and how many pieces of additional information (or assumptions and brute force) are needed to break the integrity protections. With surprises that tend to be on the side of 'Oops. Turns out this clever trick can recover the originals easier than we thought.'

> It's honestly kind of disingenuous to describe dimensionality reduction in the way that they do here. It is like reducing the resolution of a photo, but it'd best be described as reducing that resolution to say, the 20 most representative pixels. There's no real sense in which the photo still exists.

In my honest opinion the original analogy does an excellent job of intuitively explaining that most of the informative aspects of the data are kept (we can still see just fine what's in the image) while irrelevant details are discarded, and that is probably what was intended.

If anything comes off as disingenuous in that context it's your representation that it's like a strong reduction in the pixel domain (where it does indeed destroy a lot of the information). What can be done is much more like running the picture through a high-performance Imagenet classifier and keeping the 20 (or 2048, or whatever's needed) most informative values at a level that corresponds strongly to semantic content of the picture, and holding on the model. We could probably generate images that people would have a hard time distinguishing from the original with that.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#154
post #65

Earlier quoted context omitted.

Door to door canvassers these days carry devices that tell you what topics to bring up and what topics not to bring up at a certain address, even distinguishing between individuals at an address; some are told to demand a husband let them talk to the wife, for example.

I don't know about the specific campaigns that you are referring to, but in my experience a lot of the information used in campaigns I've been involved in comes from previous canvassing sessions. Political parties in most countries are involved at many levels where there are elections. Canvassing doesn't just take place for the big elections. One year they will have been round and had a lengthy discussion with Mrs X,…

Exactly. This is the old-fashioned approach to campaign targetting that Cambridge Analytica was trying (and failing) to replace: just send a bunch of volunteers to talk to them about who they're voting for and why, then put that in your big database. One of the dirty not-so-secrets about CA is that according to the Trump campaign, they were abandoned completely in favour of that old-fashioned approach because they were worse. Similarly, if you've been paying attention, you might have noticed a few insider stories about how one of the Hillary Clinton campaign's big screw-ups was underestimating the importance of that data compared to modern big data tech and basically throwing a lot of it in the trash. This didn't get nearly as much coverage as the idea that Cambridge Analytica, Trump, and Facebook were conspiring to brainwash the population, probably because it was less juicy a narrative and kind of embarassing to the Clinton campaign and the DNC.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#155
post #68
post #41

Spoiler warning. Article punchline ahead. "The whole point of a dimension reduction model is to mathematically represent the data in simpler form. It’s as if Cambridge Analytica took a very high-resolution photograph, resized it to be smaller, and then deleted the original. The photo still exists — and as long as Cambridge Analytica’s models exist, the data effectively does too." That's an eloquent piece of explanati…

> when strictly speaking the raw data has indeed been deleted after being used to create a derivative work that can for all important purposes be used to recreate the original? To be precise, you almost certainly cannot use this data to recreate anything remotely resembling the original dataset. This type of dimensionality reduction would throw away enormous volumes of data. There is no meaningful sense in which you…

This type being ... something like PCA? It's up to the user how much to actually reduce the dimensionality.

The pixel analogy is bad, but to use it anyway -- you get to choose how many pixels you keep. You could keep literally all of them.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#156

Earlier quoted context omitted.

Alternatively, "If I take a FLAC you own, make a 320kbps MP3 from it, and store it on my laptop, am I still in possession of any IP belonging to you?"

I like that analogy. I'll make it more tenuous with - "I took a copy of your album collection without your permission, ripped them to MP3, played them so much everyone is sick of them. but you've still got all the original CD's you don't even use, so no problem right?" On this tangent, IP ownership for deep learning models is interesting - how to you prove (in court) someone has/hasn't copied model/stolen a training…

Except Facebook let them take a copy of the album collection, albeit for a different use case, but it was allowed nonetheless. That doesn't absolve CA in any way, but should make us wary of people we willingly give our "album collections" - they will use them to make money, and what they allow people do with them can easily be things we don't agree with, but didn't have the imagination to think of when we signed the EULA.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#158
post #41

Spoiler warning. Article punchline ahead. "The whole point of a dimension reduction model is to mathematically represent the data in simpler form. It’s as if Cambridge Analytica took a very high-resolution photograph, resized it to be smaller, and then deleted the original. The photo still exists — and as long as Cambridge Analytica’s models exist, the data effectively does too." That's an eloquent piece of explanati…

To take a completely different approach in terms of 'derivative works':

Say we have a bunch of profile images, and then describe them in text. "Blonde, caucasian, large nose, curls, receding hairline, strong jaw, big ears", or perhaps even more specific stuff like "has a mole on the left cheek at the same height as the right earlobe" and "right nostril is larger than the left" and "dimple in chink".

Based on a description like this, we could identify an individual in probably a short paragraph. Nonetheless, on the data side, this is a lot less information than is represented by the raw pixels.

When it comes to the topic of CA's tools, and 'psychosocial' targeting, we can't separate the broader context and the way in which one single term can encode tons of data ("looks like George Clooney with a bigger forehead"). I'd argue the same princple applies to political views, and personality.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#159

Earlier quoted context omitted.

It’s kind of comparable. Regardless, I still think having the most relevant features already extracted is all they need to ask many of the questions they might want to. The point is that that’s still quite bad.

Right, I was just trying to confirm an analogy. It seems like this stuff is like a lossy codec for traits.

If you can still run Java applets, this is a nice intro: http://www.cs.mcgill.ca/~sqrt/dimr/dimreduction.html

Re: How Cambridge Analytica’s Facebook targeting model really worked

#160

Earlier quoted context omitted.

Do you have any sources for your claim about how the Clinton campaign acquired FB data and how they used it? Was any of it acquired fraudulently and/or in violation of FB's ToS, like CA's data was?

Do you have any sources for your claim that the parent poster claimed the democrats purchased Facebook data?

This is a thread about how CA acquired and used data from Facebook, so I assume the parent comment was trying to make an apples-to-apples comparison. The alternative is that the poster was disingenuously trying to imply a false equivalence.
Post reply on HN