I think the responses to this can be broken down into a 2x2 matrix: level of concern vs. understanding of technology. 1) Don't understand ML; not concerned - "I have nothing to hide." 2) Don't understand ML; concerned - "I bought this device and now people are spying on me!" 3) Understand ML; not concerned - "Of course, Google needs to label its training data." 4) Understand ML; concerned - "How can we train models/c…
Knowing what I know about how people I have worked with have come close to or have actually mishandled data despite the best of intentions, I do not trust any of these teams without an explicit accountability mechanism that is observable by an outside entity. I'm not looking to punish slip-ups, because mistakes happen, but I am looking for external enforcement to keep people honest.
It's not that I think the engineers using this data are mustache twirling villains, it's that I think mishandling is inevitable due to inattention (yes, even you make mistakes!), and we have to design our data pipelines against that.