Live data from Hacker News

Extracting audio from visual information

newsoffice.mit.edu

21–30 of 133 posts

Re: Extracting audio from visual information

#21
post #12

It doesn't seem to mention what kind of camera you'd need to do this at the mentioned distance (15 feet). I'm assuming it's at least a couple thousand bucks, or maybe even so expensive that most professional photographers won't have it, but does anyone actually know? Or did I miss it in the article?

The video goes into a bit more detail and shows a couple of things that address your question. 1) A lot of the tests were done with a camera that costs thousands of dollars. You can see images of the camera used in one of the experiments. 2) They were able to use consumer-grade cameras to also capture sound. Even frequencies up to 5x higher than the actual 60 FPS of the captured video.

I wonder how well a GoPro would do? They can do 120fps at 720p or 240fps at WVGA.

Re: Extracting audio from visual information

#22
post #7

Earlier quoted context omitted.

So it's like if 960-row video at 60fps were actually a 57600 rows-per-second video, right? Which they can extract info from because having more rows in a still frame doesn't mean having more information (at least not linearly), i.e. in still frames with no rolling shutter, rows contain redundant vibration already extracted from previous rows. So having a rolling shutter is good for this specific application because i…

Between the time the first and last row are read, the object might've moved a little bit. So if you take a picture with your phone from the side window of a moving car, the picture will appear stretched.

Sure, I meant it specifically as a guess of how it's applied to sound extraction and how it means you have ROW samples per frame.

Re: Extracting audio from visual information

#24

And now think of how much high definition video, CCTV, and other forms of recordings already exist.. and then think about running large collections of pre-existing video through such algorithms, along with the best speech-to-text in the business. You could have a whole new Wikileaks on your hands :-)

Thankfully most CCTVs are so crappy and so low framerate, I doubt they can get much out of it. If you want to record a particular person through a window, for instance, you can get a laser microphone that catches the vibrations of the glass.

Yeah, I've heard that's what the spies use, although that does require specific effort. What I find more intriguing about this development is how pre-existing footage could be used. While HD CCTV is certainly not popular, I suspect enough has been said in the presence of existing HD video to incriminate a few people :-)

Re: Extracting audio from visual information

#28

Incredible and terrifying at the same time. If they're doing this kind of stuff right now with consumer cameras, imagine how effective this technology will be in just a few decades. Privacy is fading quickly with the advent of exciting technology like this.

you need a camera that captures at least 2,000 frames per second.
Post reply on HN