Live data from Hacker News

Extracting audio from visual information

newsoffice.mit.edu

71–80 of 133 posts

Re: Extracting audio from visual information

#71
post #61
post #60

Earlier quoted context omitted.

You missed the part where the video's frames per second has to be higher than the audio frequency.

You missed the part where they proved that this is not required.

When Hollywood isn't still using 24fps film cameras they're certainly not using consumer-grade camera sensors.

Re: Extracting audio from visual information

#72
post #61
post #60

Earlier quoted context omitted.

You missed the part where the video's frames per second has to be higher than the audio frequency.

You missed the part where they proved that this is not required.

It definitely has to be 2x the desired maximum frequency if you don't want aliasing. They probably capture higher frequencies by making assumptions about the input signal, but this will inevitably lead to aliasing when these assumptions are violated. The article doesn't really go into depth about precisely how they go about this.

Re: Extracting audio from visual information

#73
post #61
post #60

Earlier quoted context omitted.

You missed the part where the video's frames per second has to be higher than the audio frequency.

You missed the part where they proved that this is not required.

High-speed camera or rolling shutter is required, (most) movies use neither.

Re: Extracting audio from visual information

#74
post #61
post #60

Earlier quoted context omitted.

You missed the part where the video's frames per second has to be higher than the audio frequency.

You missed the part where they proved that this is not required.

You missed the part where 60 frames per second only allowed to identify the speaker and the number of people in the room.

> Because of a quirk in the design of most cameras’ sensors, the researchers were able to infer information about high-frequency vibrations even from video recorded at a standard 60 frames per second. While this audio reconstruction wasn’t as faithful as it was with the high-speed camera, it may still be good enough to identify the gender of a speaker in a room; the number of speakers; and even, given accurate enough information about the acoustic properties of speakers’ voices, their identities.

Re: Extracting audio from visual information

#75
post #72
post #61

Earlier quoted context omitted.

You missed the part where they proved that this is not required.

It definitely has to be 2x the desired maximum frequency if you don't want aliasing. They probably capture higher frequencies by making assumptions about the input signal, but this will inevitably lead to aliasing when these assumptions are violated. The article doesn't really go into depth about precisely how they go about this.

They have lots of pixels, though, which means you should be able to go higher than nyquist with some clever work.

With two time samples you shouldn't be able to learn anything about the state of waves in a pool, but if each sample is a photograph with lots of pixels you can actually tell a lot.

Re: Extracting audio from visual information

#76
post #64

Now I am gonna be suspicious of chip bags and plants everywhere. We will need to invent telepathic communication (mind to mind) to preserve privacy.

> We will need to invent telepathic communication (mind to mind) to preserve privacy.

Or maybe it's time to think how to adapt to a world without privacy?

Re: Extracting audio from visual information

#77

I am just worried that they are picking up part of information from camera mic. Maybe camera mic output is not totally independent from the video sensor, and is encoded in the final video.

The camera used in most of these experiments (a Phantom high speed camera) doesn't even have a microphone - so that would be quite impossible.

Thanks. However, there is still a possibility that mechanical vibration is coming to the video sensor through some path (floor+camera tripod), and affecting the video. After all mechanical vibration is affecting the potato chips. Why is it so impossible to affect video sensor? Especially in the sensitive high speed camera? Let the downvotes begin :)

Re: Extracting audio from visual information

#78
post #17

This is from the movie Eagle Eye, right? The evil computer watches the vibrations in a cup of coffee while someone is speaking.

lol that is the first thing that I thought as well. I just passed it off as a typical Hollywood trope, but now I'm quite delighted to see it become a reality.

> I just passed it off as a typical Hollywood trope

I often wonder what makes people to pass off things like that as "typical tropes", where they are obviously realistic and doable.

Re: Extracting audio from visual information

#79
post #71
post #61

Earlier quoted context omitted.

You missed the part where they proved that this is not required.

When Hollywood isn't still using 24fps film cameras they're certainly not using consumer-grade camera sensors.

So maybe the only option would be The Hobbit?

Re: Extracting audio from visual information

#80
post #59

Wouldn't it be really neat to apply this to HD movie sequences, and hear what the sounds on the set and the voices of actors were like pre-production? And how unreal some of the sounds must have turned out with all the visual tweaking that happens in production?

I had the exact same idea.

Why did I got down voted for that? It's true. Proof: http://www.crackajack.de/2014/08/04/visual-microphone-audio-...
Post reply on HN