Live data from Hacker News

Adobe demos “photoshop for audio,” lets you edit speech as easily as text

arstechnica.com

51–58 of 58 posts

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#51
post #46
post #6

Earlier quoted context omitted.

It did sound like they picked the right things to say. I won't be surprised if the extra words appear in the 20 minutes of training speech. Either way even if it's not as magical as they try to make it its still extremely useful for the voiceover industry.

But not so much for the voice over talent.

Damn, I was just about to get into voice overs as well.

I guess I still have a couple years.

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#52
post #26

I've actually been tossing around the idea of creating a program like this, although for a specific use case. In Bethesda games (Oblivion, Skyrim, Fallout) there are large modding communities adding new quests, areas and plot lines. But one technical and financial challenge for them has always been voice acting. Not only do they have to worry about voice acting new potential characters, but they have no means of writ…

That's a very similar problem to the one in Neal Stephenson's 'Diamond Age'; the Primer that drives the story is narrated by paid 'ractors', but later the book is copied and mass-produced, and paying for all of those to be acted would be prohibitively expensive; so the voice parts are replaced by simpler, synthetic voices.

Me, I always thought the way this tech would pan out would be as voice-changing; you capture the actor's voice you want, then the writer speaks the lines and the computer restyles them as an imitation of the actor - this seemed a more natural way to input emotion and pacing.

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#53
post #26

I've actually been tossing around the idea of creating a program like this, although for a specific use case. In Bethesda games (Oblivion, Skyrim, Fallout) there are large modding communities adding new quests, areas and plot lines. But one technical and financial challenge for them has always been voice acting. Not only do they have to worry about voice acting new potential characters, but they have no means of writ…

voice acting is not strictly about reading the script

This review is a good example, Ross gives "The Chosen: Well of Souls" a _legendary bad voice acting_ badge: https://www.youtube.com/watch?v=f-5AGFf_IOg

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#54
post #28
post #16

Copied from my comment on an earlier submission on this: I don't see how the watermarking they talk about is going to succeed in preventing forgeries. If they're planning to watermark unedited recordings, you have a huge false positive problem because there are billions of hours of legitimate but unwatermarked audio recordings, and will probably continue to be. You can also get false negatives by tampering with a wat…

I doubt their watermarking can survive the analog hole and multiple reencodes.

as AnimalMuppet said under me it will survive analog hole quite well, but degrade quality. Look up Cinavia.

Its so bad you can actually hear distortions.

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#55
post #47

This sounds waaay better than the Donald Trump text to speech system I've been working on: http://jungle.horse I wish I could chat with their engineering team. I'd love to learn the mathematics and tech. (A lot of it might be patented?) Is there an equivalent of SIGGRAPH for audio?

UIST is one of the premier human-computer interaction tools conferences, so it's the closest that accepts "SIGGRAPH-like" technique papers for audio. Maneesh Agrawala from Stanford has several great papers in this space of audio mixing/editing: http://graphics.stanford.edu/~maneesh/

Top poster may also be interested in ICASSP, which bothers itself more with the algorithms of speech/audio processing than the applications of said processing.

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#56

1) Mentioned near the end of the video that it actually required around 20 minutes of audio to start synthesys. Not quite as magic as it first seemed. Still cool. 2) The intonation always matched the initial sample. Give us some filters like "vocal fry", "perplexed", "angry", "wonder" etc and then we'll really have something here.

> Give us some filters like "vocal fry", "perplexed", "angry", "wonder" etc and then we'll really have something here Neural networks can already do style transfer for images. Seems like it should be doable for voice, too.

I wonder if you could get interesting results by applying neural style transfer to spectrograms.

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#57
post #46
post #6

Earlier quoted context omitted.

It did sound like they picked the right things to say. I won't be surprised if the extra words appear in the 20 minutes of training speech. Either way even if it's not as magical as they try to make it its still extremely useful for the voiceover industry.

But not so much for the voice over talent.

Well another job that will become obsolete I guess.

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#58
post #50

You know this is a good idea when half the commenters already have a half-baked version of this created themselves!

Aha, like it was with all the pre- and post-Slack chambered forums with cloud history sync.

Did many of them become a success? Now Microsoft is in on it too, so you gotta wrestle that gorilla.

Post reply on HN