Live data from Hacker News

How Cambridge Analytica’s Facebook targeting model really worked

niemanlab.org

181–190 of 210 posts

Re: How Cambridge Analytica’s Facebook targeting model really worked

#181
post #68

Earlier quoted context omitted.

> when strictly speaking the raw data has indeed been deleted after being used to create a derivative work that can for all important purposes be used to recreate the original? To be precise, you almost certainly cannot use this data to recreate anything remotely resembling the original dataset. This type of dimensionality reduction would throw away enormous volumes of data. There is no meaningful sense in which you…

" cannot use this data to recreate anything remotely resembling the original dataset." This by itself may be mostly true perhaps - and many of the comments get into ways of playing with this dataset to make it better, I don't have experience with those methods, but, what I have not seen anyone mention, if you have this dumbed down dataset, the original is gone.. you can still combine with other data sets that are eit…

Or "better" (for some value of better), join the data to something identifiable (e.g., the public records you listed) before developing models for everything you want to retain, then discard the original data once your deep enough net has effectively auto-encoded whatever you wanted to retain in an identifiable manner.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#182
post #116

Earlier quoted context omitted.

I've heard rumours that the internal backlash at Netflix against 5 star ratings happened with Amy Schumer's most recent comedy special which had thousands of 1 star reviews on Netflix. Which was one of their most expensive comedy productions and whose release coincided not long before the switch over to vague thumbs up/down. Note, the comedy special was similarly panned across the press and social media as being repe…

By "heard rumours" do you mean "read about it on Breitbart?". http://www.breitbart.com/big-hollywood/2017/03/18/netflix-sc... Notice that they're careful to say that they made the switch "amid" the special, not because of it. Also as far as I can tell, they have no actual data on the fact, and they're the only "newspaper" reporting it.

Nope... it was a Reddit self-post (incl. people who put together the viral compilation video of her joke 'borrowing', breaking down why she's no longer as popular, and analyzing Netflix's timing and stated rationale).

I'm curious why you brought up Breitbart? I just googled it and found a ton of other (non-political/right-wing) sites which drew the same exact conclusion between Schumer and the ending of 5 star reviews. Are you trying to say her thousands of awful reviews were somehow political? Or that it's all just some "alt-right" right-wing conspiracy?

- https://movieweb.com/netflix-cancels-5-star-rating-system/

- http://ew.com/tv/2017/03/16/netflix-star-ratings/

- http://collider.com/netflix-rating-system-thumbs-up/

- http://screenertv.com/television/goodbye-stars-hello-thumbs-...

- https://www.washingtontimes.com/news/2017/mar/17/netflix-cha...

I saw Amy's show live, which she later filmed, and it was just awful. Those reviews were highly justified. And I used to be a big fan of hers before anyone knew who she was.

If Amy Schumer wasn't the original reason why they started to ditch 5-stars then it was most certainly a motivating factor to get it pushed out (conveniently right after her special was famously destroyed - which got widespread press before the switch over to thumbs). As it most certainly also faced internal resistance as they were basically abandoning a decade of 5-star reviews via user-generated-content, in favour of a vague Thumbs Up/Down.

Amy provided a perfect example of the inconvenient and conflicting goals which Netflix has as a producer and content platform. The timing of the release would have been HIGHLY coincidental if it wasn't related in some fashion.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#183
post #136
post #85

Earlier quoted context omitted.

As long as they retain no data which could specifically identify the original person, yes. There is nothing wrong with building segmentation models as long as they aren't specific enough to identify a specific person. My concern would be, how granular is too granular? What if we added "and live in zip code 12355 and is registered Green Party"? This now gets eerily specific, and might be sufficient to identify an indi…

Why would they ever discard that? Why would there be a granularity where ML suddenly stops working? Why would you even stop at one model per person, instead of one model per mood, or modes of thought at different stress points?

In fact they would desire that granularity most of all so as to reconcile the past and future state psychographic profiles for an individual- then they could attempt to isolate the causation of a state change- basically they need to identify the moment an individuals profile reflects the change from democrat to republican or vice versa. Or Religious to atheist etc.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#184
post #147
post #140

Earlier quoted context omitted.

They claim to have deleted that data. If they haven't deleted the data, then of course it's still an invasion of privacy. But the ML model really has nothing to do with it.

I think the ML model has a lot to do with it in this case. One of the arguments I expect to see is that "Oh, no! We removed all the data . It's gone. I mean, that was only a few hundred megabytes per person anyway, but we just calculate a few thousand numbers from it and save in our system, then delete the data. That's less data per person than is needed to show a short cute cat GIF. What harm could we possibly do wi…

My point isn't that there is no harm here in them storing this model. It's also not that the data in their model is worthless. It's specifically that the way this article is talking about the issue is incorrect. The analogy they use would lead you to draw false conclusions about what's going on, and how to understand it.

There is a real issue here of whether or not they should be allowed to keep a model trained from ill-gotten data. But the way I would think about it is: If you steal a million dollars and invest it in the stock market, and make a 10% return, what happens to that 10% return if you then return the original million? That's a much better analogy for what's going on here. They stole an asset, and made something from it, and it's unclear who owns that thing or what to do with it.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#185
post #90

Earlier quoted context omitted.

Except that that toy example bears no resemblance to the actual situation.

How many dimensions were they working with and how much variance and correlation was there in the features? What's the margin of error for the end product?

I don't know precisely, but it's pretty obvious that there'd be no way to reconstruct personally identifying info from it.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#186

Earlier quoted context omitted.

From this you infer that he is under the control of Putin? He is pro Russia no doubt, but I don't think that is what's being asserted by the media. I'd prefer they stick to facts, do you disagree?

He's "pro-Russia" in the sense that he seems to have some sort of admiration for Putin's tough-guy persona, but I can't see much other sense in which that's meaningfully true.

I agree, but unfortunately that minor inconvenience won't stop newspaper reporters from writing evidence free articles implying the contrary.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#187

Earlier quoted context omitted.

Personally, I used to get substantially better results. It would suggest to me a movie I'd never seen before with a score above 90, I'd watch, and enjoy. Now I'm getting things like kids shows suggested to me. I double checked by history to make sure no one watched one on my account, but they keep popping up with high percentage. I also really like horror and get weird suggestions [1] for similar movies. After I watc…

The coverage of the Netflix ratings format switch has always been frustrating to me because what people are complaining about (you're not the only one; https://www.polygon.com/2017/4/7/15212718/netflix-rating-sys... ) was predictable from the beginning. I do research in this area and it's fairly well established that when you go from something like five points to two points with ratings, you throw away tons of inform…

I mean it is pretty obvious that the loss of information is going to make it harder to predict. My big problem is that the system suggests completely off base content. Not to mention overly pushing their own content.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#188

Earlier quoted context omitted.

That's only accurate in the sense that because an LSTM's hidden layer is much smaller in dimension than the data on which it is trained, there is less information in it. However, it concisely represents a manifold in a much larger dimensional space and effectively captures most of the information in it. It may be (and is) lossy, but don't underestimate the expressive power of a deep neural network.

You're throwing out buzzwords instead of addressing the response. It's dimensionality reduction. You cannot recover the original object. It's like using a shadow to reconstruct the face of the person casting the shadow. Note this has nothing to do with the expressive power of a deep neural network. You are by definition trying to throw away noisy aspects of the data and generalize a lower dimensional manifold from a…

It's dimensionality reduction. You cannot recover the original object.

Makes me think of the Simulacrum[1]. "The map is not the territory."[2]

1. https://en.wikipedia.org/wiki/Simulacra_and_Simulation

2. https://en.wikipedia.org/wiki/Map%E2%80%93territory_relation

Re: How Cambridge Analytica’s Facebook targeting model really worked

#189
post #74

Earlier quoted context omitted.

Really? Because in terms of actual usefulness, I find youtube's suggestions to better...

I tend to agree with you. The only thing I dislike about it is when I happen to open a link I'll randomly be sent or notice in an article that's something unlike what I normally enjoy or something I close right away and then for the next few days, or until I watch a bunch more of what I usually enjoy, the entire suggestion list is only things to do with that one random link.

I don't think it's possible on a phone, but in a browser you can remove videos from your watch history. Click the three dots next to the video for options, and select not interested. Google actually seems to take these into account, I had watched an Alex Jones video that was linked from elsewhere, and while it is good to not have too much of a bubble there's just certain things I don't need to see more of.

Or, just use a separate browser. I typically only use opera for YouTube & anything where I don't mind Google tracking me, but if I open a YouTube link in Firefox I'm not logged in (and w/ opera's VPN enabled it doesn't appear to affect YouTube recommendations). I'm sure Google does correlate traffic between the two to some extent, but this seems like the only useful way to use operas integrated VPN.

Re: How Cambridge Analytica’s Facebook targeting model really worked

#190
post #152
post #92

Very interesting article but I wish it went one step further. Why does it matter that cambridge analytica knew a user's big five or that they were an old, uneducated republican? How was this (inferred) data used? I assume they wrote/created different ads for different sets of users... but how many segments did they have? Did their graphic designer build 500 different ads, or was text/images dynamically inserted based…

I wonder if you learned something since your blatant failure last week where you gave all that trust into facebooks hands, and praised advertisements, one day before they've blown up.

Oh, I forgot you were the one who called me mind rapist for working in advertising. I'm tempted to do an AMA here so people can hear the perspective of an insider (not to sway opinions necessarily, but to quell some misconceptions and understand how we got here/how things actually work)
Post reply on HN