Live data from Hacker News

Adobe demos “photoshop for audio,” lets you edit speech as easily as text

arstechnica.com

41–50 of 58 posts

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#41
post #26

I've actually been tossing around the idea of creating a program like this, although for a specific use case. In Bethesda games (Oblivion, Skyrim, Fallout) there are large modding communities adding new quests, areas and plot lines. But one technical and financial challenge for them has always been voice acting. Not only do they have to worry about voice acting new potential characters, but they have no means of writ…

https://www.youtube.com/watch?v=YPwFuCL33I8 but automated, then.

I used to love watching mans1ay3r's videos, amazing level of creativity and hard work required to assemble all the voice clips! I can only imagine what hilarity would come from a fully fledged voice-splicing application!

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#42
post #38

Joaquin Phoenix is now going to narrate all of my audio books.

I doubt this software can figure out which parts of your audiobook should be whispered, which parts emphasized, which kind of accent each character should have, when the passage should be read as a playful conversation, as opposed to a clinical narrative... A lot more goes into an audiobook then some guy reading 300 pages of text into a microphone.

I wonder how long it will be before someone gets it doing intonation correctly. I doubt it's an intractable problem and it probably wouldn't be too hard to turn existing audiobooks into a pretty decent dataset.

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#43
post #42
post #38

Earlier quoted context omitted.

I doubt this software can figure out which parts of your audiobook should be whispered, which parts emphasized, which kind of accent each character should have, when the passage should be read as a playful conversation, as opposed to a clinical narrative... A lot more goes into an audiobook then some guy reading 300 pages of text into a microphone.

I wonder how long it will be before someone gets it doing intonation correctly. I doubt it's an intractable problem and it probably wouldn't be too hard to turn existing audiobooks into a pretty decent dataset.

If you want it to not sound like a comedy of errors, you'll have to wait until we figure out Strong AI that can understand sarcasm in the works of Shakespeare... And be able to carry on a coherent, deep philosophical debate.

The difficulty is not in modulating the voice to create intonation, but in annotating prose with intonation.

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#44

Earlier quoted context omitted.

What about feeding in an audiobook narrated by a well-known personality (like Stephen Fry), and then using the voice to narrate your own self-published eBook?

Maybe this technique will lead to people "copyrighting" their voice. That would be weird word where intonation can be copyrighted.

Some people have tried to analyze this as an aspect of right of publicity (e.g. http://ir.lawnet.fordham.edu/cgi/viewcontent.cgi?article=281... but I think there are many other articles on this topic). The right of publicity analysis could be different from the copyright analysis.

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#45
post #33
post #20

Earlier quoted context omitted.

The comparison to photoshop seems a silly stretch, but > actually reconstruction a real TTS by using a small sample it actually did this in the demo, if I understand what you're asking. It created brand-new words, using the speakers own voice, from a tiny sample size. (If we believe the demo.)

At the end of the demo the speaker clarifies that it requires about 20 minutes of speech, and this was a controlled demo so it's quite possible that brand new words were actually not created.

Think about what one could do with this tool and a dataset of the speeches of a president or something. I mean, with the amount of speeches our presidents give from candidacy to finishing their term as president, we'd have probably weeks worth of words to pick from. If it's an intelligent system you could use this for so much even if it is simple cutting and pasting!

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#46
post #6
post #2

I thought it sounded badly cut when he moved wife in the sentence but adding new text was pretty amazing.

It did sound like they picked the right things to say. I won't be surprised if the extra words appear in the 20 minutes of training speech. Either way even if it's not as magical as they try to make it its still extremely useful for the voiceover industry.

But not so much for the voice over talent.

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#47
This sounds waaay better than the Donald Trump text to speech system I've been working on: http://jungle.horse

I wish I could chat with their engineering team. I'd love to learn the mathematics and tech. (A lot of it might be patented?)

Is there an equivalent of SIGGRAPH for audio?

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#48
post #47

This sounds waaay better than the Donald Trump text to speech system I've been working on: http://jungle.horse I wish I could chat with their engineering team. I'd love to learn the mathematics and tech. (A lot of it might be patented?) Is there an equivalent of SIGGRAPH for audio?

UIST is one of the premier human-computer interaction tools conferences, so it's the closest that accepts "SIGGRAPH-like" technique papers for audio. Maneesh Agrawala from Stanford has several great papers in this space of audio mixing/editing: http://graphics.stanford.edu/~maneesh/

Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text

#49

1) Mentioned near the end of the video that it actually required around 20 minutes of audio to start synthesys. Not quite as magic as it first seemed. Still cool. 2) The intonation always matched the initial sample. Give us some filters like "vocal fry", "perplexed", "angry", "wonder" etc and then we'll really have something here.

It would be really useful to capture speed, intonation and other things involving expression from a sample from somebody else. It seems quite hard to describe those as text is not enough, and even harder to correctly generate them from a database that contains no matching expression. I imagine that currently you can't turn the speech into a track for a rollercoaster scream.

[deleted]
Post reply on HN