Live data from Hacker News

Extracting audio from visual information

newsoffice.mit.edu

11–20 of 133 posts

Re: Extracting audio from visual information

#11

And now think of how much high definition video, CCTV, and other forms of recordings already exist.. and then think about running large collections of pre-existing video through such algorithms, along with the best speech-to-text in the business. You could have a whole new Wikileaks on your hands :-)

[deleted]

Re: Extracting audio from visual information

#12
It doesn't seem to mention what kind of camera you'd need to do this at the mentioned distance (15 feet). I'm assuming it's at least a couple thousand bucks, or maybe even so expensive that most professional photographers won't have it, but does anyone actually know? Or did I miss it in the article?

Re: Extracting audio from visual information

#13

And now think of how much high definition video, CCTV, and other forms of recordings already exist.. and then think about running large collections of pre-existing video through such algorithms, along with the best speech-to-text in the business. You could have a whole new Wikileaks on your hands :-)

Thankfully most CCTVs are so crappy and so low framerate, I doubt they can get much out of it.

If you want to record a particular person through a window, for instance, you can get a laser microphone that catches the vibrations of the glass.

Re: Extracting audio from visual information

#15
post #12

It doesn't seem to mention what kind of camera you'd need to do this at the mentioned distance (15 feet). I'm assuming it's at least a couple thousand bucks, or maybe even so expensive that most professional photographers won't have it, but does anyone actually know? Or did I miss it in the article?

it says about the frame rate necessary, we can only infer the price of the camera.

Re: Extracting audio from visual information

#16
This is amazing. Imagine taking high definition video of a crowd of people. You could pick out objects nearby and hear what individuals are saying. You could nearly produce a 3D auditorium by sampling different points in a video. Couple this with a 3D camera and an Oculus Rift, you could have something incredible.

Re: Extracting audio from visual information

#18
post #7

Earlier quoted context omitted.

Cameras basically read their sensors one row of pixels at a time. By measuring the distortion of each row, they can detect vibrations higher than the camera's frame rate.

So it's like if 960-row video at 60fps were actually a 57600 rows-per-second video, right? Which they can extract info from because having more rows in a still frame doesn't mean having more information (at least not linearly), i.e. in still frames with no rolling shutter, rows contain redundant vibration already extracted from previous rows. So having a rolling shutter is good for this specific application because i…

Between the time the first and last row are read, the object might've moved a little bit. So if you take a picture with your phone from the side window of a moving car, the picture will appear stretched.

Re: Extracting audio from visual information

#19
post #12

It doesn't seem to mention what kind of camera you'd need to do this at the mentioned distance (15 feet). I'm assuming it's at least a couple thousand bucks, or maybe even so expensive that most professional photographers won't have it, but does anyone actually know? Or did I miss it in the article?

The video goes into a bit more detail and shows a couple of things that address your question.

1) A lot of the tests were done with a camera that costs thousands of dollars. You can see images of the camera used in one of the experiments.

2) They were able to use consumer-grade cameras to also capture sound. Even frequencies up to 5x higher than the actual 60 FPS of the captured video.

Post reply on HN