Live data from Hacker News

How Cambridge Analytica’s Facebook targeting model really worked

niemanlab.org

121–130 of 210 posts

Re: How Cambridge Analytica’s Facebook targeting model really worked

#121

Earlier quoted context omitted.

[ Deleted. Nothing I say on HN ever matters. Move along ]

Very few "undecided" voters truly are; elections are won and lost by getting your supporters to go to the polls. So if you wanted to use scurrilous, fake news to help your candidate, you'd be better off sending stories that will get your supporters really fired up and eager to vote and get their friends to vote, not trying to persuade the practically nonexistent undecided demographic.

Are you able to specify the sources based on which you are supporting these claims, namely that elections are not not decided by "undecided" voters but rather by pushing your supporters to the election polls?

Re: How Cambridge Analytica’s Facebook targeting model really worked

#122

I'm excited for GDPR.. The hard part is going to be getting the truth out of these companies about the actual extent of the data they hold on us

That "hard part" is going to be more than just hard; I think the word you're looking for there is "impossible". They don't have any way of knowing what data Google, Facebook, Amazon, or any other company has. As this article does a credible job of explaining, you have to understand a fair amount about statistics (PCA) and machine learning to even know if if you were looking at it, and they won't know where to look fo…

Sure. But at least when they are actually whistleblowed (by people who "just" work for the company) there is a law which can be used to call to a court the theople who are legally accountable for the company

Re: How Cambridge Analytica’s Facebook targeting model really worked

#123
post #120

Earlier quoted context omitted.

You're right that it's really just a form of dimensionality reduction. My point was just that it's a more powerful form of dimensionality reduction than PCA or NMDS. [Edit: and that the salient characteristics are likely contained in the model.]

Precisely because it's more powerful, it doesn't encode the identifying information of the original data. Something like PCA likely would retain identifying characteristics (depending on how many low-rank vectors you drop).

Outside of the fact that they have identities for all of the people whose data they acquired, yes, it would be harder to reconstruct individual people with it than PCA because of the direct interpretability of its data.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#124

Earlier quoted context omitted.

There doesn't have to be a security breach for it to be a very bad example of using data collected in one way for a completely different purpose. It violates the 'lawful basis for processing' part of privacy legislation.

Which US legislation?

That app collected data on many more than just US residents so more than just US legislation applies. This is one of those pesky little problems of doing stuff 'on the internet', especially when you start doing stuff that is purposefully or accidentally illegal.

Besides that they apparently also used similar trickery in their consultancy for the Brexit side.

https://www.theguardian.com/politics/2018/mar/26/pressure-gr...

This is far from over.

https://www.theguardian.com/commentisfree/2018/mar/23/plenty...

Re: How Cambridge Analytica’s Facebook targeting model really worked

#125

Earlier quoted context omitted.

You're throwing out buzzwords instead of addressing the response. It's dimensionality reduction. You cannot recover the original object. It's like using a shadow to reconstruct the face of the person casting the shadow. Note this has nothing to do with the expressive power of a deep neural network. You are by definition trying to throw away noisy aspects of the data and generalize a lower dimensional manifold from a…

You're right that it's really just a form of dimensionality reduction. My point was just that it's a more powerful form of dimensionality reduction than PCA or NMDS. [Edit: and that the salient characteristics are likely contained in the model.]

Is this basically a choice between .mp3 and .ogg, png vs jpg vs gif?

Re: How Cambridge Analytica’s Facebook targeting model really worked

#126

Earlier quoted context omitted.

Can you point to the part of that article containing evidence of to what degree they affected the election?

The admission by the company executive.

That's interesting. How would he know the degree to which he influenced the election? Believing any claims to somehow fact rather than plain old self-promotion seems rather naive, or am I missing something?

Re: How Cambridge Analytica’s Facebook targeting model really worked

#127
post #83
post #80

Earlier quoted context omitted.

How do you know what CA trained on, or what's possible? Do you have qualifications in ML?

I know what they trained on because it's been reported on. They got around 50 million people's FB profiles, and a smaller subset's (300k, I think) personality test results. I use ML models every day in my work, and understand how they function. It is true that individuals information is probabilistically encoded into the parameters of the model. However, if the model is any good, the people they trained on's informat…

Is there a reason that people are only talking about the privacy angle?

People very much don't want these models to exist. They don't want a predictive model which will guess their affiliation just by providing unrelated Activity bread crumbs.

That's why I assumed this whole issue has exploded recently.

Not the privacy, but the implications.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#128

Earlier quoted context omitted.

You're right that it's really just a form of dimensionality reduction. My point was just that it's a more powerful form of dimensionality reduction than PCA or NMDS. [Edit: and that the salient characteristics are likely contained in the model.]

Is this basically a choice between .mp3 and .ogg, png vs jpg vs gif?

It’s kind of comparable.

Regardless, I still think having the most relevant features already extracted is all they need to ask many of the questions they might want to. The point is that that’s still quite bad.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#129
post #68
post #41

Spoiler warning. Article punchline ahead. "The whole point of a dimension reduction model is to mathematically represent the data in simpler form. It’s as if Cambridge Analytica took a very high-resolution photograph, resized it to be smaller, and then deleted the original. The photo still exists — and as long as Cambridge Analytica’s models exist, the data effectively does too." That's an eloquent piece of explanati…

> when strictly speaking the raw data has indeed been deleted after being used to create a derivative work that can for all important purposes be used to recreate the original? To be precise, you almost certainly cannot use this data to recreate anything remotely resembling the original dataset. This type of dimensionality reduction would throw away enormous volumes of data. There is no meaningful sense in which you…

> distill some insights about people from this data

i'd argue that insight is the bit that's important, and the bit that's the privacy risk.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#130

Earlier quoted context omitted.

Very few "undecided" voters truly are; elections are won and lost by getting your supporters to go to the polls. So if you wanted to use scurrilous, fake news to help your candidate, you'd be better off sending stories that will get your supporters really fired up and eager to vote and get their friends to vote, not trying to persuade the practically nonexistent undecided demographic.

Are you able to specify the sources based on which you are supporting these claims, namely that elections are not not decided by "undecided" voters but rather by pushing your supporters to the election polls?

I thought this was common enough knowledge not to want citations, but I think you will find these satisfactory.

http://www.stat.columbia.edu/~gelman/research/unpublished/sw...

https://www.politico.com/magazine/story/2014/01/independent-...

https://www.thenation.com/article/what-everyone-gets-wrong-a...

The last piece has a short summary of the salient point:

> In fact, according to an analysis of voting patterns conducted by Michigan State University political scientist Corwin Smidt, those who identify as independents today are more stable in their support for one or the other party than were “strong partisans” back in the 1970s. According to Dan Hopkins, a professor of government at the University of Pennsylvania, “independents who lean toward the Democrats are less likely to back GOP candidates than are weak Democrats.”

> While most independents vote like partisans, on average they’re slightly more likely to just stay home in November. “Typically independents are less active and less engaged in politics than are strong partisans,” says Smidt.

> [...]

> The conventional wisdom holds that the parties need independents to win general elections, but the reality is that they’re increasingly devoting their resources to getting their own voters—including their “closet partisans”—out to the polls rather than trying to sway the dwindling number of genuine swing voters. “We’ve seen a huge increase in technology and the ability to turn out the vote,” says Smidt. “So in terms of a cost-benefit analysis, the parties and candidates see that it’s much easier to turn out people who agree with them than it is to change someone’s mind. And then there’s also the question of how many of us are even open to changing our minds.”

Post reply on HN