Live data from Hacker News

Extracting audio from visual information

newsoffice.mit.edu

91–100 of 133 posts

Re: Extracting audio from visual information

#91
post #83

Earlier quoted context omitted.

Thanks. However, there is still a possibility that mechanical vibration is coming to the video sensor through some path (floor+camera tripod), and affecting the video. After all mechanical vibration is affecting the potato chips. Why is it so impossible to affect video sensor? Especially in the sensitive high speed camera? Let the downvotes begin :)

I fail to see how that would make this any less impressive.

If my evil :) speculation is true, the camera sensor is just picking up mechanical vibrations from environment (coming through air, tripod etc.) and encoding them into video. There is no proof that the bag of chips is vibrating (they even say that the vibrations are not visible on footage). They are extracting something from video, but that something may just be spuriuos pickups by equipment, and not related image/video of bag of chips. Thus it would be just a silly way to measure spurious mechanical vibrations.

Re: Extracting audio from visual information

#92
post #51

Earlier quoted context omitted.

There is some small possibility of improvement through software techniques, such as maybe data assimilation, which can use information from surrounding time-frames to improve the measurement. This is assuming that the magnitude of vibrations changes a lot slower than the vibrations themselves, which is usually true, and how most audio compression works. It may be able to clean up the sound a little. However, I would…

The data comes in faster than 60 fps. A camera sensor doesn't capture the entire frame instantly every 1/60 second. It progressively scans through the frame over some measurable fraction of that 1/60 second. This is that quirk. Suppose the camera scans 720 lines in HD every 1/60 second. Each row is offset in time by 1/43200 second. A rigid object could be slightly offset in space on each line of pixels, indicating th…

Well the reader would read as fast as it can.

Let's say that it would read the entire image in 1/120 second, then it is waiting and does nothing another 1/120 second before it starts reading next frame.

The real number would be significantly smaller. Therefore they can not bump the sample rate more then five or six times. And I imagine they are using some intelligent algorithm to evenly space out the captured samples already.

Re: Extracting audio from visual information

#93

While it's cool to see this technique applied to consumer video, the general idea has been used in international espionage for a while now. In fact, Léon Theremin (yes, of the musical instrument) invented an early device capable of eavesdropping based on window vibrations: http://en.wikipedia.org/wiki/L%C3%A9on_Theremin#Espionage . I also recall that this is why all of the windows in the White House are fitted with t…

Well this system would in theory bypass the White House countermeasures

Re: Extracting audio from visual information

#94
post #61

Earlier quoted context omitted.

You missed the part where they proved that this is not required.

You missed the part where 60 frames per second only allowed to identify the speaker and the number of people in the room. > Because of a quirk in the design of most cameras’ sensors, the researchers were able to infer information about high-frequency vibrations even from video recorded at a standard 60 frames per second. While this audio reconstruction wasn’t as faithful as it was with the high-speed camera, it may s…

If you watch the video in the post, the demo is surprisingly accurate.

Re: Extracting audio from visual information

#96

While it's cool to see this technique applied to consumer video, the general idea has been used in international espionage for a while now. In fact, Léon Theremin (yes, of the musical instrument) invented an early device capable of eavesdropping based on window vibrations: http://en.wikipedia.org/wiki/L%C3%A9on_Theremin#Espionage . I also recall that this is why all of the windows in the White House are fitted with t…

The technologies involved are quite different and this new approach radically simplifies the logistics.

Re: Extracting audio from visual information

#97
How much vibration does sound actually cause an object. I tried humming near various objects and noticed nothing. I used to be really interested in early mechanical microphones and sound recording, but I couldn't find much information on how they work either.

Re: Extracting audio from visual information

#98
post #79
post #71

Earlier quoted context omitted.

When Hollywood isn't still using 24fps film cameras they're certainly not using consumer-grade camera sensors.

So maybe the only option would be The Hobbit?

The Hobbit still most likely doesn't use rolling shutter sensors, but I'm willing to be proven wrong.

Re: Extracting audio from visual information

#99

Wouldn't it be really neat to apply this to HD movie sequences, and hear what the sounds on the set and the voices of actors were like pre-production? And how unreal some of the sounds must have turned out with all the visual tweaking that happens in production?

I would expect that you need to acquire the raw footage; post processing (and compression) will ruin the data.

Re: Extracting audio from visual information

#100

Wouldn't it be really neat to apply this to HD movie sequences, and hear what the sounds on the set and the voices of actors were like pre-production? And how unreal some of the sounds must have turned out with all the visual tweaking that happens in production?

I record the dialog for your movies. Aprt from obvious sound effects like Darth Vader voices or so, actors' dialog is changed as little as possible during post production. It's not heavily EQed or comrpessed, because that would mean a corresponding change to the room tone and background noise, which would then fluctuate unnaturally as you went back and forth between the participants in a scene. In scenes with a lot of movement actors' voices sometimes sound a little deeper than in real life, because they are fitted with a tiny wireless microphone, and there's usually some resonance from the chest cavity.

But generally what you hear is very close to how the person actually sounds - although their accent or inflections may be adopted for the purposes of their role. This can be a bit jarring; I've worked with method actors who maintain their screen accent at all times during production until the film is done, so when they switch back to their regular accent after a month or so it's extremely disorienting, since I've been listening to them in my headphones day in day out for weeks, and am paid to pay as much attention to their voices as the cinematographer pays to their faces.

Some actors go even farther in support of their public image. Rock Hudson had a somewhat high voice that producers deemed incompatible with his looks, so during production he would warm up every day by shouting for 20 minutes and gargling with orange juice to inflame his vocal cords, and of course he smoked a lot too. What actors will do to themselves in pursuit of screen presence far exceeds anything I've ever been asked to do in post. Editing dialog is more than enough work without trying to sculpt people's voices.

Post reply on HN