I think the responses to this can be broken down into a 2x2 matrix: level of concern vs. understanding of technology. 1) Don't understand ML; not concerned - "I have nothing to hide." 2) Don't understand ML; concerned - "I bought this device and now people are spying on me!" 3) Understand ML; not concerned - "Of course, Google needs to label its training data." 4) Understand ML; concerned - "How can we train models/c…
If the training already only happens on a 1/500 sample, skewing the sample towards "people who don't care about their privacy" will probably not significantly impact the quality of the data.
I'm surprised this wasn't already the case, but hopefully the article will help the people responsible make better decisions in the trade-off between minimizing onboarding friction and respecting user's privacy in the future.