Live data from Hacker News

How Cambridge Analytica’s Facebook targeting model really worked

niemanlab.org

91–100 of 210 posts

Re: How Cambridge Analytica’s Facebook targeting model really worked

#91
post #87
post #83

Earlier quoted context omitted.

I know what they trained on because it's been reported on. They got around 50 million people's FB profiles, and a smaller subset's (300k, I think) personality test results. I use ML models every day in my work, and understand how they function. It is true that individuals information is probabilistically encoded into the parameters of the model. However, if the model is any good, the people they trained on's informat…

But if the resulting model doesn't contain information about individuals, how does this help targeting individuals for the campaign? Edit: is it that the model is then applied to only strictly public data about the person? If so I guess the interesting question then becomes whether the model is definitely not anything near overfitting (i.e. containing enough information to match a person's public data directly since…

> But if the resulting model doesn't contain information about individuals, how does this help targeting individuals for the campaign?

I don't know exactly what they were modeling, but from the published reports, it sounds like they were trying to predict big 5 personality characteristics (conscientousness, neuroticism, openness, extraversion, agreeableness) from FB profile data (e.g. likes, dislikes, bio, post content, etc.). So in that case, the model would contain weights that measure the strength of relationship between characteristics like "likes punk rock music" and "openness". That description really only literally applies to a linear model - but nonlinear models are, for these purposes, the same.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#92
Very interesting article but I wish it went one step further. Why does it matter that cambridge analytica knew a user's big five or that they were an old, uneducated republican? How was this (inferred) data used?

I assume they wrote/created different ads for different sets of users... but how many segments did they have? Did their graphic designer build 500 different ads, or was text/images dynamically inserted based on these variables? How did they figure out which message would resonate with each segment? How did they test something like this, with so many potential variables? Was this knowledge used only on facebook, or across all digital channels? Was it implemented in non-digital channels as well?

I'd kill to have access to their campaign set ups.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#93
post #40

Earlier quoted context omitted.

Who said anything about a security breach? Most of the controversy has been about the company influencing elections using data scraped from people (and their friends) unaware of what the data was being used for.

The degree to which it influenced the election is questionable. Despite all the headlines, I haven't yet seen any convincing analysis of the impact of facebook on the election (I'm not sure how one would even go about doing so). So far it seems like it's just a convenient vehicle for people that dislike the outcome of the election to express indignation.

Don't be naive

http://www.bbc.com/news/world-43476762

Re: How Cambridge Analytica’s Facebook targeting model really worked

#94
post #71
post #68

Earlier quoted context omitted.

> when strictly speaking the raw data has indeed been deleted after being used to create a derivative work that can for all important purposes be used to recreate the original? To be precise, you almost certainly cannot use this data to recreate anything remotely resembling the original dataset. This type of dimensionality reduction would throw away enormous volumes of data. There is no meaningful sense in which you…

> It's arguable whether they should be allowed to keep those insights, but there's no privacy risk there really. So if Google has distilled someone's emails over the years into "closeted homosexual with a deeply repressed leather fetish", that's not an invasion of their privacy as long as they throw away the source materials?

Since the source data was deleted - according to current standards and policies - their hands are probably technically clean. But there may be another angle of attack.

In the US, you're not allowed to benefit directly from a crime you committed. For example, if you rob a bank, you can't buy your mother a car with the money and say "sorry, it's gone!" when the police come knocking.

With that line of reasoning and if there was a legal, privacy, or at least a TOS breach in collecting the data, the derivative machine learning models may be tainted also. Then again, it's likely impossible to prove exactly what data went into the model, so hard to establish which models might be tainted.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#95

Earlier quoted context omitted.

That's only accurate in the sense that because an LSTM's hidden layer is much smaller in dimension than the data on which it is trained, there is less information in it. However, it concisely represents a manifold in a much larger dimensional space and effectively captures most of the information in it. It may be (and is) lossy, but don't underestimate the expressive power of a deep neural network.

Only if the "true" data actually lives in a lower dimensional manifold and the data acurrately can encode it with low noise. I doubt anyone can tell who you will vote for depending on which cat videos you liked, no matter how magic your regressor.

I do think that the most significant components of a personality will likely be targetable with a relatively low nuclear norm.

And, for example, where someone's proclivity on the exploration/exploitation spectrum, if you will, (IE, how strongly do they respond to fear-based messaging) falls is probably quite predictable from a spectrum of likes.

Cat pictures may be less informative, but not all of these people clicked exclusively on feline fuzzy photos.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#96
post #68
post #41

Spoiler warning. Article punchline ahead. "The whole point of a dimension reduction model is to mathematically represent the data in simpler form. It’s as if Cambridge Analytica took a very high-resolution photograph, resized it to be smaller, and then deleted the original. The photo still exists — and as long as Cambridge Analytica’s models exist, the data effectively does too." That's an eloquent piece of explanati…

> when strictly speaking the raw data has indeed been deleted after being used to create a derivative work that can for all important purposes be used to recreate the original? To be precise, you almost certainly cannot use this data to recreate anything remotely resembling the original dataset. This type of dimensionality reduction would throw away enormous volumes of data. There is no meaningful sense in which you…

[deleted]

Re: How Cambridge Analytica’s Facebook targeting model really worked

#97

When talk about CA first emerged on HN before the election some posters found the original papers referred to. They were looking at pictures in the story and zoomed in to find the titles. I cannot find those posts for the life of me again. Not suggesting anything nefarious here, I just can't find them. Does anyone have a link to those early conversations or make copies of the papers? I made copies earlier but deleted…

Here are some leads:

https://news.ycombinator.com/item?id=14486365

https://news.ycombinator.com/item?id=14393991

https://news.ycombinator.com/item?id=14330547

https://news.ycombinator.com/item?id=14284502

https://news.ycombinator.com/item?id=13939814

And the query: https://hn.algolia.com/?query=mercer&sort=byPopularity&prefi...

Re: How Cambridge Analytica’s Facebook targeting model really worked

#98

Earlier quoted context omitted.

Only if the "true" data actually lives in a lower dimensional manifold and the data acurrately can encode it with low noise. I doubt anyone can tell who you will vote for depending on which cat videos you liked, no matter how magic your regressor.

I do think that the most significant components of a personality will likely be targetable with a relatively low nuclear norm. And, for example, where someone's proclivity on the exploration/exploitation spectrum, if you will, (IE, how strongly do they respond to fear-based messaging) falls is probably quite predictable from a spectrum of likes. Cat pictures may be less informative, but not all of these people clicke…

I'm not an expert on personality so won't disagree (except to say that I am a little sceptical of a static personality profile actually existing and I think people who always vote a certain way would be the easiest to regress and also the most useless to target). As I said in another post, it really depends on what part of your privacy you are trying to protect. It is also a mistake to think of anything on the internet as a private forum.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#99
post #87
post #83

Earlier quoted context omitted.

I know what they trained on because it's been reported on. They got around 50 million people's FB profiles, and a smaller subset's (300k, I think) personality test results. I use ML models every day in my work, and understand how they function. It is true that individuals information is probabilistically encoded into the parameters of the model. However, if the model is any good, the people they trained on's informat…

But if the resulting model doesn't contain information about individuals, how does this help targeting individuals for the campaign? Edit: is it that the model is then applied to only strictly public data about the person? If so I guess the interesting question then becomes whether the model is definitely not anything near overfitting (i.e. containing enough information to match a person's public data directly since…

[deleted]

Re: How Cambridge Analytica’s Facebook targeting model really worked

#100
post #75
post #71

Earlier quoted context omitted.

> It's arguable whether they should be allowed to keep those insights, but there's no privacy risk there really. So if Google has distilled someone's emails over the years into "closeted homosexual with a deeply repressed leather fetish", that's not an invasion of their privacy as long as they throw away the source materials?

If they kept information like that, then yes that would be an invasion of privacy. But that sort of information is almost certainly not encoded in an ML model trained on 50 million people's data.

Let's say I take age and income of everyone in a city and train a regression model that predicts income from age. The model has slope and intercept that "encode" the information from all the people.

It would not be possible to make inferences about the income of any particular person from the slope and intercept, so it would be ok to share those values in, say, a journal article, even though disclosing income of a particular person would not be ok.

Post reply on HN