Live data from Hacker News

Algorithm recovers speech from a potato-chip bag filmed through glass (2014)

news.mit.edu

91–100 of 111 posts

Re: Algorithm recovers speech from a potato-chip bag filmed through glass (2014)

#92
post #60
post #45

Earlier quoted context omitted.

> [Jones] claimed to have extracted the hum of the potter's wheel from the grooves of a pot, So far so good... > and the word "blue" from an analysis of patch of blue color in a painting. What the hell?

I don’t know about the techniques but the theory isn’t prima facie impossible. A paintbrush can act like a microphone just like anything else, and if the paintbrush (more likely a putty knife or more rigid object) picked up a sound while applying paint, that could manifest in the paint layer.

>> I don’t know about the techniques but the theory isn’t prima facie impossible.

The Total Perspective Vortex derives its picture of the whole Universe on the principle of extrapolated matter analyses.To explain — since every piece of matter in the Universe is in some way affected by every other piece of matter in the Universe, it is in theory possible to extrapolate the whole of creation — every sun, every planet, their orbits, their composition and their economic and social history from, say, one small piece of fairy cake.

- Douglas Adams

Re: Algorithm recovers speech from a potato-chip bag filmed through glass (2014)

#93
post #15

At some point it is (controversially) hypothesized that we may be able to pull imprinted recordings off of ancient artifacts: https://en.wikipedia.org/wiki/Archaeoacoustics#Past_interpre...

Anyone remember the TV show Fringe? I seem to remember them using science fiction tech like that. Life imitates art.

It was capturing telephone touch tones on a plate of glass.

Olivia's smartphone could dial based on touch tones, that's totally sci fi these days.

Re: Algorithm recovers speech from a potato-chip bag filmed through glass (2014)

#94

Earlier quoted context omitted.

With enough processing power, you can listen to anyone as long as you can get an object in frame that’s vibrating with speech (the UK and Chinese CCTV networks comes to mind). Depending on resolution, I’d expect this to be the case for any footage already stored. It’s just a matter of having a distributed computing job kicked off to comb through video data and add the additional audio metadata. Will you only speak of…

Don't sweat the CCTV or recordings: Reconstructing audio from video requires that the frequency of the video samples — the number of frames of video captured per second — be higher than the frequency of the audio signal. In some of their experiments, the researchers used a high-speed camera that captured 2,000 to 6,000 frames per second. That’s much faster than the 60 frames per second possible with some smartphones,…

The article quite literally is about a technique to bypass that limitation and do it with a regular 60fps camera.

Re: Algorithm recovers speech from a potato-chip bag filmed through glass (2014)

#95
"If I'm sitting next to a swimming pool, and somebody dives in - and she's not too pretty, so I can think of something else - I think of the waves and things that have formed in the water. And, uh, when there's lots of people have dived in the pool there's a very great choppiness of all these waves all over the water and to think that it's possible, maybe, that in those waves there's a clue as to what's happening in the pool. That some sort of insect or something with sufficient cleverness could sit in the corner of the pool and just be disturbed by the waves, and by the nature of the irregularities and bumping of the waves have figured out who jumped in where and when and where what's happening all over the pool. And that's what we're doing when we're looking at something. Uh, the light that comes out is ... is waves, just like in the swimming pool except in three dimensions instead of the two dimensions of the pool it's they're going in all directions. And we have a eighth of an inch black hole into which these things go ... which, uh, is particularly sensitive to the parts of the waves that are coming in a particular direction it's not particularly sensitive when they're coming in at the wrong angle which we say is from the corner of our eye. And if we want to get more information from the corner of our eye we swivel this ball about so that the hole moves from place to place. Then ... uh, it's quite wonderful that we can see ... figure out so easy. That's really because the light waves are easier than the ... the waves in the water are a little bit more complicated it would have been harder for the bug than for us but it's the same idea. Figure out what the thing is that we're looking at at a distance."

Transcribed from footage included in the documentary "The Last Journey of a Genius" (1989) by Christopher Sykes, a BBC TV production in association with WGBH Boston and Coronet/MTI Film and Video.

https://www.youtube.com/watch?v=1qQQXTMih1A

Re: Algorithm recovers speech from a potato-chip bag filmed through glass (2014)

#96

Earlier quoted context omitted.

Don't sweat the CCTV or recordings: Reconstructing audio from video requires that the frequency of the video samples — the number of frames of video captured per second — be higher than the frequency of the audio signal. In some of their experiments, the researchers used a high-speed camera that captured 2,000 to 6,000 frames per second. That’s much faster than the 60 frames per second possible with some smartphones,…

The article quite literally is about a technique to bypass that limitation and do it with a regular 60fps camera.

It sounds like that technique doesn't quite recover the audio though, just a portion of it:

While this audio reconstruction wasn’t as faithful as that with the high-speed camera, it may still be good enough to identify the gender of a speaker in a room; the number of speakers; and even, given accurate enough information about the acoustic properties of speakers’ voices, their identities.

Saying it may still be good enough pretty strongly implies that the quality is quite low.

Re: Algorithm recovers speech from a potato-chip bag filmed through glass (2014)

#97
post #63
post #61

Earlier quoted context omitted.

And...someone was saying “blue” the instant they applied the blue paint? That’s certainly not impossible to believe, it just seems like a bit of a stretch.

Maybe it was painted by Bob Ross?

Titanium white.

Re: Algorithm recovers speech from a potato-chip bag filmed through glass (2014)

#98

Earlier quoted context omitted.

The article quite literally is about a technique to bypass that limitation and do it with a regular 60fps camera.

It sounds like that technique doesn't quite recover the audio though, just a portion of it: While this audio reconstruction wasn’t as faithful as that with the high-speed camera, it may still be good enough to identify the gender of a speaker in a room; the number of speakers; and even, given accurate enough information about the acoustic properties of speakers’ voices, their identities. Saying it may still be good e…

There is a video on the webpage that includes the recovered audio. It is low quality, but enough to understand the speaker most of the time, and definately enough to pose a security risk when analysed by a specialist

Re: Algorithm recovers speech from a potato-chip bag filmed through glass (2014)

#99

Forgive me, as someone with zero actual knowledge in this field. This may be a naive question, but wouldn't this be fairly easily defeated by playing songs in the background?

Thanks for the replies everyone. A follow up: would the sound recovered from such vibrations have enough resolution to perform some of the tricks described (isolating voice from background noise or identifying the song and cancelling it out, etc...)?

Re: Algorithm recovers speech from a potato-chip bag filmed through glass (2014)

#100

Earlier quoted context omitted.

It sounds like that technique doesn't quite recover the audio though, just a portion of it: While this audio reconstruction wasn’t as faithful as that with the high-speed camera, it may still be good enough to identify the gender of a speaker in a room; the number of speakers; and even, given accurate enough information about the acoustic properties of speakers’ voices, their identities. Saying it may still be good e…

There is a video on the webpage that includes the recovered audio. It is low quality, but enough to understand the speaker most of the time, and definately enough to pose a security risk when analysed by a specialist

The youtube clip in the article? The voice reconstruction there is from high speed video. If you mean some other video, I'd appreciate a link.
Post reply on HN