Live data from Hacker News

Adobe demos “photoshop for audio,” lets you edit speech as easily as text

arstechnica.com

11–20 of 58 posts

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#11
post #4

Earlier quoted context omitted.

Photos are still admissible, despite Photoshop existing for several decades.

People were convincingly doctoring photos looong before Photoshop.

Communist countries are notorious for making people disappear.

http://www.businessinsider.com/people-who-were-erased-from-h...

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#12
I wonder how much better the rendering would be if the audio track were much longer and the software would have more to learn. I don't mean more words for a 1 to 1 match since it's clearly beyond that (pronouncing words that it didn't see), but voice features that weren't in that short track.

Hypothetical question: say it had access to all the episodes from Key & Peele, would the rendering be better to the point you could basically generate an audio track from a script with intonation and all?

It would be interesting if they offered "voice packages" either online or offline so you could just pass text through it and the output would be a Morgan Freeman narration. You'd have a shop for "Cords" the same as iTunes for songs & apps. Maybe game developers will find that interesting, too. Having access to way more voices than they'd have in real life, and on a budget.

Someone could also save their voices for posterity. Many people listen to recordings of loved ones who passed away to remember them. Saving the voice for new content would be something to think about.

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#13
post #4

Earlier quoted context omitted.

Photos are still admissible, despite Photoshop existing for several decades.

To be fair on that front: all evidence can be challenged as being false/fake/etc

Right, a photo (or recording) in court is only as good as what you can prove about its history. Same as with any piece of evidence, really.

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#14
If photoshop would give these results, a lot of industries would go belly up.

Nice marketing line, but it's speech recognition which set the begin/end frame in the sample.

I was expecting either "painting" away defects or actually reconstruction a real TTS by using a small sample.

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#16
Copied from my comment on an earlier submission on this:

I don't see how the watermarking they talk about is going to succeed in preventing forgeries.

If they're planning to watermark unedited recordings, you have a huge false positive problem because there are billions of hours of legitimate but unwatermarked audio recordings, and will probably continue to be. You can also get false negatives by tampering with a watermark-capable device to get it to watermark something that wasn't recorded from analog. Or you can rerecord edited audio from an analog source and simply claim that your "genuine" recording is slightly noisy.

If they're planning to watermark edited recordings, someone else can implement the same kind of technology but without the watermarking.

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#17

How long before video and audio evidence will not longer be admissible in court? 2030?

This election, Facebook was full of fake news stories. The next Presidential election, Facebook is going to be full of fake video clips.

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#20

If photoshop would give these results, a lot of industries would go belly up. Nice marketing line, but it's speech recognition which set the begin/end frame in the sample. I was expecting either "painting" away defects or actually reconstruction a real TTS by using a small sample.

The comparison to photoshop seems a silly stretch, but

> actually reconstruction a real TTS by using a small sample

it actually did this in the demo, if I understand what you're asking. It created brand-new words, using the speakers own voice, from a tiny sample size. (If we believe the demo.)

Post reply on HN