Live data from Hacker News

Facebook engineers: we have no idea where we keep all your personal data

theintercept.com

131–137 of 137 posts

Re: Facebook engineers: we have no idea where we keep all your personal data

#131

This is misdirection plain and simple. In the best most convincing way. Get tunnel vision (sorry) nerds to make an honest display of their ignorance of the bigger picture at Meta. Aww shucks we don't know where that data comes from.... Of course senior engineers don't know where the data comes from exactly. Does your mechanic know where each tire comes from or care? Does your fav restaurant chef know each field that…

> there are most definitively some people that know at least roughly "At least roughly" is doing a lot of work here. I think what you're getting at is that there are people who know which storage systems are likely to hold that data, but knowing exactly which objects do is far more challenging. Copies get made while data pipelines are tweaked, and then forgotten. Copies get made internally on storage systems as disks…

I'm saying they have vast mountains of data on people. They make enough money from it to be one of the most profitable companies on the planet. This is not happening by accident.

I'm quite certain that if I showed up to FB with a $50 million dollar ad campaign to run they would have quite the opposite story about their data knowledge. Marching dozens of people into the room promising they can slice and dice every bit of info down to the atom for my needs.

Yet in court they want to claim they live by the seat of their pants. Its a deception tactic. Even though I'm sure its as you describe partly true in technical concept.

I do believe that FB is the least privacy invasive of all the big tech giants for what its worth. Nothing against them. I actually think this case is an intentional aggravation to the courts to have regulation created on data collection. That way FAANGM can cement themselves as the offical national data brokers with costly startup regulation for any compeditors.

Re: Facebook engineers: we have no idea where we keep all your personal data

#132
post #11

In many ways you can think of large, long-living tech companies not unlike old cities like, say, London or Paris. The buildings and roads you see are built on top of older buildings and ruins. The streets are weirdly shaped and intersect at odd angles because they were made hundreds of years before and adapted over time as needs evolved. There are catacombs underneath sidewalks and no one genuinely understands it all…

> The buildings and roads you see are built on top of older buildings and ruins

It's easy to delete old data that is never accessed, so I don't get your point...

Re: Facebook engineers: we have no idea where we keep all your personal data

#133

Earlier quoted context omitted.

> there are most definitively some people that know at least roughly "At least roughly" is doing a lot of work here. I think what you're getting at is that there are people who know which storage systems are likely to hold that data, but knowing exactly which objects do is far more challenging. Copies get made while data pipelines are tweaked, and then forgotten. Copies get made internally on storage systems as disks…

I'm saying they have vast mountains of data on people. They make enough money from it to be one of the most profitable companies on the planet. This is not happening by accident. I'm quite certain that if I showed up to FB with a $50 million dollar ad campaign to run they would have quite the opposite story about their data knowledge. Marching dozens of people into the room promising they can slice and dice every bit…

> I'm quite certain

Based on what?

> they would have quite the opposite story about their data knowledge

As I said, they can probably delete everything they could sell. What's left are fragments, which are still a real concern but not enough to use for ad targeting and certainly not enough to satisfy a big-money customer. They might take that $50M anyway, but it would be under false pretenses and that's a very different issue.

Re: Facebook engineers: we have no idea where we keep all your personal data

#134
I would not be surprised at all. I can count on one hand the number of Engineering candidates that given a system description involving three visible networked laptops in view of the camera have answered "where is the data, right now?" , or can at least give a decent stab at "trace the datapath through the machine".

I can't explain somehow this is the average candidate, yet somehow... Life goes on.

Re: Facebook engineers: we have no idea where we keep all your personal data

#135
Not surprising that the same can be said for a certain US government department having hosted a personal mail server in a bathroom closet and trafficking classified materials and caught much later by OIG.

Or that one time I found a dialup modem under the raised floor and attached to the Equifax (then TRW) mainframe.

Or catching someone inserting a USB stick into a PC inside a secured white-lab area.

Absolutely crazy times for large entities.

Re: Facebook engineers: we have no idea where we keep all your personal data

#136
post #114

Earlier quoted context omitted.

as per my other comment, in most of the sufficiently big or historical cases you have to distinguish between two things: - reality - the socially acceptable fiction in the report the exercise would be to create an output that allows people to pretend we know where all instances of "log4j" were used, or that sufficient depth and resources have been expended on such. but in a sufficiently large and complex organisation…

On a spectrum of ridiculousness vs. necessity running from "enumerate where cutlery was used for the past one hundred years" to "keep tabs on the radioactive material we have in storage", it is a choice to vacuum up user data and treat it as the former rather than the latter. It is not unavoidable; it is that Facebook does not want to avoid it. If they did want to avoid it, they could do so in the same way that my em…

That was an impressively long threaded argument just to agree with what the original poster stated. Essentially this boils down to the following: cities and large tech companies are hugely complicated endeavors and if enough money and/or manpower is placed into a "documentation project" we could fully describe either system (using a combination of reality and socially acceptable fiction.) However, it would mostly likely be outdated by the time the project was finished - unless we froze the city/company in place which would most likely kill it.

Re: Facebook engineers: we have no idea where we keep all your personal data

#137
post #114

Earlier quoted context omitted.

On a spectrum of ridiculousness vs. necessity running from "enumerate where cutlery was used for the past one hundred years" to "keep tabs on the radioactive material we have in storage", it is a choice to vacuum up user data and treat it as the former rather than the latter. It is not unavoidable; it is that Facebook does not want to avoid it. If they did want to avoid it, they could do so in the same way that my em…

That was an impressively long threaded argument just to agree with what the original poster stated. Essentially this boils down to the following: cities and large tech companies are hugely complicated endeavors and if enough money and/or manpower is placed into a "documentation project" we could fully describe either system (using a combination of reality and socially acceptable fiction.) However, it would mostly lik…

No, I do not agree with what the original poster stated.

I think it's possible to get too abstract in discussion with an idea of "documenting a system" since that can of course shoot for a level of detail that can't be maintained. Here is a practical example.

https://en.wikipedia.org/wiki/Personal_Information_Protectio...

When this law was to come into effect, every company that does business in China had to audit where they were storing all Chinese PII in order to figure out what they might need to change. For large companies, this involved making everyone perform this audit for their own little fiefdoms.

Being as this wasn't the kind of thing one can have lawyers go challenge in court, there were two options:

1. Actually do an audit (and if you're not dealing with exceptionally poor engineering, this should not have been as hard as keeping track of dependency libraries), rebuild everything from scratch (wasteful), or nuke absolutely anything that might have a chance of having Chinese PII (...unlikely from a business perspective).

2. Lie (internally / externally) and thereby break the law.

People in these comments seem to love the phrase "socially acceptable fiction" and I want to be very crisp that we are dealing with a clean dichotomy. You do the audit, or you lie and break the law.

Yes, I understand there are some among us who have so little integrity that they'll shrug at the latter option, that they'd fake the VW emissions test, that they'll cheat and embezzle and tell themselves that it was all inevitable anyway and anyone in their places would do the same. But this is not actually true. We would not all do this. Pretending this is normal variance gives cover to scummy people.

I therefore object to the root comment's characterization of this kind of thing exceeding real understanding as a natural consequence of scale/time. Yes, in terms of how it can come about that no individual knows this stuff, it's accurate – but only in the absence of the institution itself caring to pay attention. I object to succeeding characterizations that it is an inevitable attribute of social systems that meaningful control is impossible, because even educated and carefully hired professionals can apparently only be expected to behave with the ethics one might expect of a sixth grader who would rather let Chegg write his book report.

From the article, this is what FB is saying:

> “We do not have an adequate level of control and explainability over how our systems use data, and thus we can’t confidently make controlled policy changes or external commitments such as ‘we will not use X data for Y purpose,’” the 2021 document read.

This may be accurate about Facebook; I would not know. What I do know is that it is not an inevitable consequence of scale and lifespan.

Post reply on HN