Live data from Hacker News

Extracting audio from visual information

newsoffice.mit.edu

1–10 of 133 posts

Re: Extracting audio from visual information

#4

Because of a quirk in the design of most cameras’ sensors, the researchers were able to infer information about high-frequency vibrations even from video recorded at a standard 60 frames per second. Can anyone explain what this quirk is?

It's explained further on in the press release. The "quirk" is the same one that causes rolling shutter artifacts in videos: reading the camera's sensor one row at a time, instead of all at once.

Re: Extracting audio from visual information

#5

Because of a quirk in the design of most cameras’ sensors, the researchers were able to infer information about high-frequency vibrations even from video recorded at a standard 60 frames per second. Can anyone explain what this quirk is?

Cameras basically read their sensors one row of pixels at a time. By measuring the distortion of each row, they can detect vibrations higher than the camera's frame rate.

Re: Extracting audio from visual information

#7

Because of a quirk in the design of most cameras’ sensors, the researchers were able to infer information about high-frequency vibrations even from video recorded at a standard 60 frames per second. Can anyone explain what this quirk is?

Cameras basically read their sensors one row of pixels at a time. By measuring the distortion of each row, they can detect vibrations higher than the camera's frame rate.

So it's like if 960-row video at 60fps were actually a 57600 rows-per-second video, right? Which they can extract info from because having more rows in a still frame doesn't mean having more information (at least not linearly), i.e. in still frames with no rolling shutter, rows contain redundant vibration already extracted from previous rows.

So having a rolling shutter is good for this specific application because it trades off resolution (most of which is redundant or insignificant information) for sampling rate.

Re: Extracting audio from visual information

#8
And now think of how much high definition video, CCTV, and other forms of recordings already exist.. and then think about running large collections of pre-existing video through such algorithms, along with the best speech-to-text in the business. You could have a whole new Wikileaks on your hands :-)

Re: Extracting audio from visual information

#10
post #9

Am I the only one who isn't getting any audio from the video at all? I tried in two different browsers, downloaded the video with youtube-dl and tried to play it with mpv, everything to no avail, it's just a video with no sound.

Pretty ironic ... I wonder if they could run their algorithm on the video on their page ...
Post reply on HN