Live data from Hacker News

Adobe demos “photoshop for audio,” lets you edit speech as easily as text

arstechnica.com

21–30 of 58 posts

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#21

I wonder how much better the rendering would be if the audio track were much longer and the software would have more to learn. I don't mean more words for a 1 to 1 match since it's clearly beyond that (pronouncing words that it didn't see), but voice features that weren't in that short track. Hypothetical question: say it had access to all the episodes from Key & Peele, would the rendering be better to the point you…

> It would be interesting if they offered "voice packages" either online or offline so you could just pass text through it and the output would be a Morgan Freeman narration.

On that note, should Morgan Freeman be allowed to copyright the sound of his own voice?

My immediate inclination is No, you shouldn't be allowed to copyright a voice timbre because it opens up so many weird possibilities for making even more of a mess of copyright law. For example, what happens if two people happen to have the same voice, or at least, close enough so that Adobe's software can't distinguish them.

But there is a reasonable case to be made that Morgan Freeman's voice belongs to him. If some company profits from having a reproduction of Morgan Freeman's voice in one of their products, maybe they should pay him.

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#22

I wonder how much better the rendering would be if the audio track were much longer and the software would have more to learn. I don't mean more words for a 1 to 1 match since it's clearly beyond that (pronouncing words that it didn't see), but voice features that weren't in that short track. Hypothetical question: say it had access to all the episodes from Key & Peele, would the rendering be better to the point you…

One of the things that seems limiting with this version is that, when the words existed (like "wife" and "dogs"), it just seemed to copy-and-paste them around. That honestly sounded pretty bad. But the new words sounded better.

I think that if they were using some of the AI-based libraries for generating speech, like WaveNet[1] (hear the samples towards the bottom), and then layered on-top the learning that this algorithm did to get it to sound like the speaker, you'd have a pretty powerful system that would be able to get the intonation right much of the time.

And then, of course, when you have the Pro version, you're presumably be able to add additional markup to specify intonation even further.

1. https://deepmind.com/blog/wavenet-generative-model-raw-audio...

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#23
post #21

I wonder how much better the rendering would be if the audio track were much longer and the software would have more to learn. I don't mean more words for a 1 to 1 match since it's clearly beyond that (pronouncing words that it didn't see), but voice features that weren't in that short track. Hypothetical question: say it had access to all the episodes from Key & Peele, would the rendering be better to the point you…

> It would be interesting if they offered "voice packages" either online or offline so you could just pass text through it and the output would be a Morgan Freeman narration. On that note, should Morgan Freeman be allowed to copyright the sound of his own voice? My immediate inclination is No, you shouldn't be allowed to copyright a voice timbre because it opens up so many weird possibilities for making even more of…

Morgan Freeman's likeness - including his voice - is already protected legally, no?

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#24
if adobe has this working in a demo, rest assured "security service" developed such thing 10 years ago. then you can go back and ask yourselves why osama has been reported dead as early as like 2001, the cia released videos in which he always looked different and why his body was quickly drowned at an unknown location.

go back to sleep, now. everything's alright. great new tech. will help catching terrorists from beneath your bed.

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#25
post #24

if adobe has this working in a demo, rest assured "security service" developed such thing 10 years ago. then you can go back and ask yourselves why osama has been reported dead as early as like 2001, the cia released videos in which he always looked different and why his body was quickly drowned at an unknown location. go back to sleep, now. everything's alright. great new tech. will help catching terrorists from ben…

Wait. America made up Osama. Killed him in 2001. But quietly so no one would know. Then digitally created new videos of him, all to kill him in 2011. Uh huh

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#26
I've actually been tossing around the idea of creating a program like this, although for a specific use case.

In Bethesda games (Oblivion, Skyrim, Fallout) there are large modding communities adding new quests, areas and plot lines. But one technical and financial challenge for them has always been voice acting. Not only do they have to worry about voice acting new potential characters, but they have no means of writing new dialogue for existing characters.

In Fallout 4, for example, the protagonist is fully voice acted. That means a distinct change between the way the main game feels, and any modding efforts made by the community (barring actually re-hiring the original voice actor for new lines).

I'm envisioning having this tool train on the already provided voice lines in the game(depending on the character in question, that's quite a bit). And then letting mod authors input new dialogue lines to be spit out in somewhat the actors voice.

Lots of problems with the approach of course, not to mention the fact that these are actors and not just voices (there would probably be significant amount of emotion lost). But it would give the modding community such a powerful tool to add new plots for existing characters.

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#28
post #16

Copied from my comment on an earlier submission on this: I don't see how the watermarking they talk about is going to succeed in preventing forgeries. If they're planning to watermark unedited recordings, you have a huge false positive problem because there are billions of hours of legitimate but unwatermarked audio recordings, and will probably continue to be. You can also get false negatives by tampering with a wat…

I doubt their watermarking can survive the analog hole and multiple reencodes.

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#29
post #28
post #16

Copied from my comment on an earlier submission on this: I don't see how the watermarking they talk about is going to succeed in preventing forgeries. If they're planning to watermark unedited recordings, you have a huge false positive problem because there are billions of hours of legitimate but unwatermarked audio recordings, and will probably continue to be. You can also get false negatives by tampering with a wat…

I doubt their watermarking can survive the analog hole and multiple reencodes.

I doubt that much audio quality can survive that, either.

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#30
post #26

I've actually been tossing around the idea of creating a program like this, although for a specific use case. In Bethesda games (Oblivion, Skyrim, Fallout) there are large modding communities adding new quests, areas and plot lines. But one technical and financial challenge for them has always been voice acting. Not only do they have to worry about voice acting new potential characters, but they have no means of writ…

https://www.youtube.com/watch?v=YPwFuCL33I8 but automated, then.
Post reply on HN