"None worked farther than 5 feet away" Well, that seems pretty easy to combat.
Hackers send silent commands to speech recognition systems with ultrasound
51–60 of 88 posts
Re: Hackers send silent commands to speech recognition systems with ultrasound
#52As the article says, there is a physical presence bar to meet to make this workable now, but I can't not think about clever ways that this could become a "worm" or spread digitally. Infected viral (actually viral!) YouTube videos? Robocalls that you hope go on speakerphone (where the fidelity of the attack signal may be questionable)? Or drive a car with big speakers around a neighborhood? How about a mobile phone bo…
Re: Hackers send silent commands to speech recognition systems with ultrasound
#53This is MUCH bigger deal than most understand. This will cost less than $10 to build and their is no hardware solution on phones or Alexia. Phreaking is back.
Unplug Alexa, turn "Hey Siri" off. Schlep to store for Cheerios.
Re: Hackers send silent commands to speech recognition systems with ultrasound
#54These speech recognition tools need to have some sort of authentication: - How about having a secret "wake word" instead of "Alexa" or "Hey Siri"? - Only treating signals using human voice range - Voice identification If this isn't patched soon (excluding ultrasounds), it could mean that these tools are already using inaudible signals for other purposes. For example, commercials could add ultrasounds to know who's wa…
> Only treating signals using human voice range Inter-modulation will let you create something hardware can't tell isn't. Two inaudible sounds, both getting received end up looking like an audible one to the hardware. The harmonic effect described in the article is similar, and damn hard to filter against physically. > For example, commercials could add ultrasounds to know who's watching them Yep, and we're already t…
Re: Hackers send silent commands to speech recognition systems with ultrasound
#55Earlier quoted context omitted.
Yeah, those parts are easily available on Aliexpress in droves. Imagine this in the crowded subway. "hey siri" "show me pictures of CENSORED " "send the first picture to mom" "yes, send it" :S
I’m pretty sure that Hey Siri is keyed to the users voice to some extent, so Siri should be safe. Not sure about Google, Cortana, Alexa etc.
Re: Hackers send silent commands to speech recognition systems with ultrasound
#56Re: Hackers send silent commands to speech recognition systems with ultrasound
#57Re: Hackers send silent commands to speech recognition systems with ultrasound
#58Earlier quoted context omitted.
Yeah, those parts are easily available on Aliexpress in droves. Imagine this in the crowded subway. "hey siri" "show me pictures of CENSORED " "send the first picture to mom" "yes, send it" :S
I’m pretty sure that Hey Siri is keyed to the users voice to some extent, so Siri should be safe. Not sure about Google, Cortana, Alexa etc.
I take that as, Siri is probably not trained to only the wake-up phrase, and as such, seems to be _not_ trained to the voice at all since this attack worked.
That, or poorly trained. Don't know, can't verify.
Re: Hackers send silent commands to speech recognition systems with ultrasound
#59Re: Hackers send silent commands to speech recognition systems with ultrasound
#60Earlier quoted context omitted.
> The idea is that if you want to create a frequency of "A", you can emit two powerful tones at frequencies "B" and "B+A", where the frequency B is high enough to be out of hearing range. The non-linearity of the microphone means the two tones mix together to produce a number of other frequencies, including the frequency "B-A"-"B" = "A". Does this work for ears, too? If so, are the non-linearities of different people…
Yes it does: https://makezine.com/2008/10/08/homebrew-parametric-speak/ http://www.soundlazer.com/ It's reasonably consistent. Differences in non-linearity will result in different amplitudes for each intermodulation product, but not different frequencies. Typically these systems use the "third order" product. I gather that the non-linearity exploited is as much a property of the air as the ear.
If someone is listening to a live musical instrument that is producing both audible sound and ultrasonic sound [1], is what the person perceives affected by intermodulation in the ear?
If the performance is also recorded using a technology that for all practical purposes reproduces perfectly everything in the audible range, then I can see a couple possible cases.
1. The microphone is designed to filter out ultrasonics or is sufficient linear to not have intermodulation.
In this case, the recording is what would be heard with no intermodulation. When played back, all the listener gets is the audible portion of the original sound, without any ultrasonics. Thus there is nothing to produce intermodulation in the listner's ear, and so the listener might perceive the recording as having a different timbre than the live instrument.
2. The microphone does not filter ultrasonics and is non-linear enough to have intermodulation. The audible intermodulation products will then be included in the recording.
When played back the listener will hear intermodulation products, but they will be the ones from the microphone's non-linearity, not the ear's non-linearity.
The question then is how close are microphone non-linearities to ear non-linearities. If they are similar, then the timbre of the recording should match live. If they are sufficiently different, the timbre could sound off.
It should be possible to design a system that records only audible frequencies and plays back only audible frequencies and sounds identical to live, but it may require specifically taking into account ultrasonics instead of just cutting them out like I think we currently do.
[1] A trumpet with a Harmon mute playing a quiet note has about 2% of its energy above 20 KHz. Playing a loud note drops that to about 0.5%. A cymbal crash is about 40% above 20 KHz. (Keys jangling are almost 70% above 20 KHz, which probably has something to do with why back in the early days of TV remote controls when they were ultrasonic instead of IR or RF people would report that if someone's keys jangled the channel would sometimes change). See: https://www.cco.caltech.edu/~boyk/spectra/spectra.htm