Live data from Hacker News

Adobe demos “photoshop for audio,” lets you edit speech as easily as text

arstechnica.com

31–40 of 58 posts

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#31
post #21

Earlier quoted context omitted.

> It would be interesting if they offered "voice packages" either online or offline so you could just pass text through it and the output would be a Morgan Freeman narration. On that note, should Morgan Freeman be allowed to copyright the sound of his own voice? My immediate inclination is No, you shouldn't be allowed to copyright a voice timbre because it opens up so many weird possibilities for making even more of…

Morgan Freeman's likeness - including his voice - is already protected legally, no?

If you're using "audio with substantial similarity" to the voice of Morgan Freeman, but using it to create entirely new works... doesn't that create new and interesting collisions between rights of publicity, fair use, transformative works...?

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#32
1) Mentioned near the end of the video that it actually required around 20 minutes of audio to start synthesys. Not quite as magic as it first seemed. Still cool.

2) The intonation always matched the initial sample. Give us some filters like "vocal fry", "perplexed", "angry", "wonder" etc and then we'll really have something here.

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#33
post #20

If photoshop would give these results, a lot of industries would go belly up. Nice marketing line, but it's speech recognition which set the begin/end frame in the sample. I was expecting either "painting" away defects or actually reconstruction a real TTS by using a small sample.

The comparison to photoshop seems a silly stretch, but > actually reconstruction a real TTS by using a small sample it actually did this in the demo, if I understand what you're asking. It created brand-new words, using the speakers own voice, from a tiny sample size. (If we believe the demo.)

At the end of the demo the speaker clarifies that it requires about 20 minutes of speech, and this was a controlled demo so it's quite possible that brand new words were actually not created.

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#34

1) Mentioned near the end of the video that it actually required around 20 minutes of audio to start synthesys. Not quite as magic as it first seemed. Still cool. 2) The intonation always matched the initial sample. Give us some filters like "vocal fry", "perplexed", "angry", "wonder" etc and then we'll really have something here.

It would be really useful to capture speed, intonation and other things involving expression from a sample from somebody else. It seems quite hard to describe those as text is not enough, and even harder to correctly generate them from a database that contains no matching expression. I imagine that currently you can't turn the speech into a track for a rollercoaster scream.

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#35

How long before video and audio evidence will not longer be admissible in court? 2030?

It won't be inadmissable but there should be continued push-back to establish authenticity. This is not a new problem – e.g. Nikon has sold things like http://imaging.nikon.com/lineup/software/img_auth/ for years – but it's definitely far from complete.

In most cases, this is going to come back to third-party services – think about the history of things like mailing a copy of a document to yourself using registered mail or having it notarized so the timestamp would be admissable in court. I'd be surprised if we haven't already seen someone submit timestamps from something like Facebook in court but we could use a general notary-as-a-service for this kind of app as more and more of the world moves online.

That won't help with a determined fraud from the beginning but that's always been a challenge in the court system and so most of that will come down to continuing to break down the attitude that digital data is somehow immune to the same trust issues as every other human artifact.

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#36

1) Mentioned near the end of the video that it actually required around 20 minutes of audio to start synthesys. Not quite as magic as it first seemed. Still cool. 2) The intonation always matched the initial sample. Give us some filters like "vocal fry", "perplexed", "angry", "wonder" etc and then we'll really have something here.

> Give us some filters like "vocal fry", "perplexed", "angry", "wonder" etc and then we'll really have something here

Neural networks can already do style transfer for images. Seems like it should be doable for voice, too.

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#37

Just imagine what this will do for dubbing anime or any other tv show. It's still scary how this can be abused.

What about feeding in an audiobook narrated by a well-known personality (like Stephen Fry), and then using the voice to narrate your own self-published eBook?

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#38

Joaquin Phoenix is now going to narrate all of my audio books.

I doubt this software can figure out which parts of your audiobook should be whispered, which parts emphasized, which kind of accent each character should have, when the passage should be read as a playful conversation, as opposed to a clinical narrative...

A lot more goes into an audiobook then some guy reading 300 pages of text into a microphone.

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#39

Just imagine what this will do for dubbing anime or any other tv show. It's still scary how this can be abused.

What about feeding in an audiobook narrated by a well-known personality (like Stephen Fry), and then using the voice to narrate your own self-published eBook?

Maybe this technique will lead to people "copyrighting" their voice. That would be weird word where intonation can be copyrighted.

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#40
post #4

How long before video and audio evidence will not longer be admissible in court? 2030?

Photos are still admissible, despite Photoshop existing for several decades.

Documents are still admissible, even though there are copy machines known to alter documents when copying in certain cases:

http://www.dkriesel.com/en/blog/2013/0802_xerox-workcentres_...

Post reply on HN