I've actually been tossing around the idea of creating a program like this, although for a specific use case. In Bethesda games (Oblivion, Skyrim, Fallout) there are large modding communities adding new quests, areas and plot lines. But one technical and financial challenge for them has always been voice acting. Not only do they have to worry about voice acting new potential characters, but they have no means of writ…
https://www.youtube.com/watch?v=YPwFuCL33I8 but automated, then.
Adobe demos “photoshop for audio,” lets you edit speech as easily as text
41–50 of 58 posts
Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text
#42Joaquin Phoenix is now going to narrate all of my audio books.
I doubt this software can figure out which parts of your audiobook should be whispered, which parts emphasized, which kind of accent each character should have, when the passage should be read as a playful conversation, as opposed to a clinical narrative... A lot more goes into an audiobook then some guy reading 300 pages of text into a microphone.
Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text
#43Earlier quoted context omitted.
I doubt this software can figure out which parts of your audiobook should be whispered, which parts emphasized, which kind of accent each character should have, when the passage should be read as a playful conversation, as opposed to a clinical narrative... A lot more goes into an audiobook then some guy reading 300 pages of text into a microphone.
I wonder how long it will be before someone gets it doing intonation correctly. I doubt it's an intractable problem and it probably wouldn't be too hard to turn existing audiobooks into a pretty decent dataset.
The difficulty is not in modulating the voice to create intonation, but in annotating prose with intonation.
Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text
#44Earlier quoted context omitted.
What about feeding in an audiobook narrated by a well-known personality (like Stephen Fry), and then using the voice to narrate your own self-published eBook?
Maybe this technique will lead to people "copyrighting" their voice. That would be weird word where intonation can be copyrighted.
Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text
#45Earlier quoted context omitted.
The comparison to photoshop seems a silly stretch, but > actually reconstruction a real TTS by using a small sample it actually did this in the demo, if I understand what you're asking. It created brand-new words, using the speakers own voice, from a tiny sample size. (If we believe the demo.)
At the end of the demo the speaker clarifies that it requires about 20 minutes of speech, and this was a controlled demo so it's quite possible that brand new words were actually not created.
Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text
#46I thought it sounded badly cut when he moved wife in the sentence but adding new text was pretty amazing.
It did sound like they picked the right things to say. I won't be surprised if the extra words appear in the 20 minutes of training speech. Either way even if it's not as magical as they try to make it its still extremely useful for the voiceover industry.
Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text
#47I wish I could chat with their engineering team. I'd love to learn the mathematics and tech. (A lot of it might be patented?)
Is there an equivalent of SIGGRAPH for audio?
Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text
#48This sounds waaay better than the Donald Trump text to speech system I've been working on: http://jungle.horse I wish I could chat with their engineering team. I'd love to learn the mathematics and tech. (A lot of it might be patented?) Is there an equivalent of SIGGRAPH for audio?
Re: Adobe demos “photoshop for audio,” lets you edit speech as easily as text
#491) Mentioned near the end of the video that it actually required around 20 minutes of audio to start synthesys. Not quite as magic as it first seemed. Still cool. 2) The intonation always matched the initial sample. Give us some filters like "vocal fry", "perplexed", "angry", "wonder" etc and then we'll really have something here.
It would be really useful to capture speed, intonation and other things involving expression from a sample from somebody else. It seems quite hard to describe those as text is not enough, and even harder to correctly generate them from a database that contains no matching expression. I imagine that currently you can't turn the speech into a track for a rollercoaster scream.