I mean. Of course they are. Do you expect to be able to do any meaningful level of training on data that hasn't been properly labeled? At some point, a human has to go in and correct the software when the software gets it wrong. If you want services that do what Google Home does, you have to have this. Even with that, I'm sure that the engineers are flagging voice requests that happen more then once, or where some on…
How about Google PAY MONEY to generate training data or gather it from informed people? (“Make $100 by recording your voice for Google’s machine learning algorithms.”) not “Arms full with an infant? This $50 device will solve all your problems” and then shipping those recordings to third party contractors in unsecured facilities.
Google employees are listening to Google Home conversations
171–180 of 463 posts
Re: Google employees are listening to Google Home conversations
#172Earlier quoted context omitted.
Aldous Huxley got there earlier with A Brave New World as well. Speculative Fiction was ahead of the curve in that era.
Anyone want to recommend a book? Let me join in by suggesting the Asimov robot books, which I’ve just begun
Also check out The Culture novels by Iain M Banks (he was the closest recent author to be a great the equal of Wells, Huxley or Orwell (imo) just a phenomenal writer).
The Polity series by Neal Asher are well imagined and he world builds brilliantly.
Re: Google employees are listening to Google Home conversations
#173Earlier quoted context omitted.
This classification is very useful to discuss this issue. The difference between 3 and 4, noble as it is, can be caused by feasability concerns that push people into 3, not just ignorance of the privacy impact. Human labelling of training data sets is a big thing in supervised learning. Methods that dispense with this would be valuable for purely economic reasons beyond privacy - the cost of human labelling of data s…
It's not wrong for humans to label training data. It's wrong to let humans listen to voice recordings that users believed would be between them and a computer. The solutions are obvious: sell the things with a big sticker that says, "don't say anything private in earshot," revert to old fashioned research methods where you pay people to participate in your studies and get their permission, or ask people for permissio…
Re: Google employees are listening to Google Home conversations
#174So Google’s response is (paraphrased as fairly as I can while removing the sugar-coating): ’Yes, we hire people to listen in to and transcribe some conversations from the private homes of our customers (so as improve our speech recognition engines); but the recordings aren’t linked to personally identifiable information.’ Even assuming they have only the purest intentions here, I still don’t understand how they can p…
I would count a recording of my voice as "personally identifiable information" right off the bat. Voice printing is a thing, and anyone will also tell you that they recognize the voices of people they interact with regularly. If someone played an audio clip of someone I know talking to Google Assistant to me, I would recognize who it was based on their voice.
Authentication on fixed phrases is reasonably accurate within a very few words, so at minimum it should be possible to associate "Hey Google" clips with regular users of Google Assistant voice control (i.e. "OK Google"). Identifying whether someone is present in a large dataset on open phrases is much harder, but a ~30s clip could do the job fairly consistently for anyone with access to a significant amount of voice data. And if this employee (who isn't directly working for Google) shared 'thousands' of clips with a news org, the cautious bet is that some other employee might share them with anyone willing to pay for the records.
Re: Google employees are listening to Google Home conversations
#175Earlier quoted context omitted.
How about Google PAY MONEY to generate training data or gather it from informed people? (“Make $100 by recording your voice for Google’s machine learning algorithms.”) not “Arms full with an infant? This $50 device will solve all your problems” and then shipping those recordings to third party contractors in unsecured facilities.
Maybe if you get you training data by some very different source than your real data comes from, it won't be representative and won't work very well?
I cannot find the blog post now, but quite a few years ago I recall some Google employees noticed a large number of queries for "cha cha cha cha cha..." from Android users in New York. All of the queries were done using voice search, so they listened to a few of the recordings. It turns out that their speech-to-text models were interpreting the sound of the NYC metro pulling into a station as speech.
Obviously they didn't have enough training data of people trying to talk next to a train.
Re: Google employees are listening to Google Home conversations
#176I think the responses to this can be broken down into a 2x2 matrix: level of concern vs. understanding of technology. 1) Don't understand ML; not concerned - "I have nothing to hide." 2) Don't understand ML; concerned - "I bought this device and now people are spying on me!" 3) Understand ML; not concerned - "Of course, Google needs to label its training data." 4) Understand ML; concerned - "How can we train models/c…
5) Understand ML; concerned - "Why do other people in the ML industry think it's OK to use and store peoples data in without informed consent (which are only those in group 3, and group 1+2 don't have informed consent)"
Re: Google employees are listening to Google Home conversations
#177Earlier quoted context omitted.
I mean, it's roughly the same level of exposure you run into literally every minute you're out in public anyway - all those phones around you have microphones, and more and more, they're always listening in exactly the same fashion. This is how always-on "Hey Siri" or "Okay Google" works. Which isn't to say it isn't frustrating, just that your frustration at your friend is misplaced - they're not exposing you to anyt…
> they're always listening in exactly the same fashion. This is how always-on "Hey Siri" or "Okay Google" works Afaik it is not, actually. The devices are taught to recognize the wakeup words (like "ok google") completely offline, and only the stuff recorded afterwards gets uploaded to the cloud of contractors.
Re: Google employees are listening to Google Home conversations
#178"What Orwell failed to predict is that we'd buy the cameras ourselves, and that our biggest fear would be that nobody was watching."
I think Frank Herbert hit it a bit more on the nose: “Once men turned their thinking over to machines in the hope that this would set them free. But that only permitted other men with machines to enslave them.” Frank captures the motivation, the fact that it is desirable is the scary part. Akin to the genetic engineering scene from Gattaca.
Re: Google employees are listening to Google Home conversations
#179I think the responses to this can be broken down into a 2x2 matrix: level of concern vs. understanding of technology. 1) Don't understand ML; not concerned - "I have nothing to hide." 2) Don't understand ML; concerned - "I bought this device and now people are spying on me!" 3) Understand ML; not concerned - "Of course, Google needs to label its training data." 4) Understand ML; concerned - "How can we train models/c…
I am sure we can whip up some such AI easily with some quantum computing. Preferably in a blockchain so we can verify correct operation and scale better.
ducks
Re: Google employees are listening to Google Home conversations
#180I think the responses to this can be broken down into a 2x2 matrix: level of concern vs. understanding of technology. 1) Don't understand ML; not concerned - "I have nothing to hide." 2) Don't understand ML; concerned - "I bought this device and now people are spying on me!" 3) Understand ML; not concerned - "Of course, Google needs to label its training data." 4) Understand ML; concerned - "How can we train models/c…
I understand ML and I'm concerned, but my position isn't "how can we train models/collect data in an ethical way?" Its more like "Do we really need to train models like this?" I continually come back to the same thought: The Amish (to take a random example) consume probably thousands of times less than I do. They do without the conveniences of technology which I "enjoy". They have neither computers nor an endless str…
Love it or hate it, our current world would not have developed the way it has if everyone was following the Amish way of life. They enjoy a massive number of benefits that were born out of alternative-to-them life styles. For example we may have "an endless stream of distracting media" but we also have a hugely increased life span and reduced mortality rate due to advances in modern science.
To me, this is a bit like the "self-made" billionaire ignoring all the societal infrastructure afforded to them..