Predicting Expressive Speaking Style from Text in End-To-End Speech Synthesis
1–10 of 10 posts
Re: Predicting Expressive Speaking Style from Text in End-To-End Speech Synthesis
#2The style tokens result in pretty incredible and realistic audio.
Re: Predicting Expressive Speaking Style from Text in End-To-End Speech Synthesis
#3Wow, these audio samples are incredible. I'm surprised to hear the model actually outputting natural-sounding breathing between and inside sentences. Most TTS systems explicitly remove things like that, but the addition of breathing makes it sound so much more natural. The style tokens result in pretty incredible and realistic audio.
Re: Predicting Expressive Speaking Style from Text in End-To-End Speech Synthesis
#4Re: Predicting Expressive Speaking Style from Text in End-To-End Speech Synthesis
#5Re: Predicting Expressive Speaking Style from Text in End-To-End Speech Synthesis
#6This seems like it could be great for automatically generating audio books. Personally I would one day like to have a program that can read arbitrary text to me in a more or less human way, that would allow me to read papers for work while driving.
Re: Predicting Expressive Speaking Style from Text in End-To-End Speech Synthesis
#7This seems like it could be great for automatically generating audio books. Personally I would one day like to have a program that can read arbitrary text to me in a more or less human way, that would allow me to read papers for work while driving.
Re: Predicting Expressive Speaking Style from Text in End-To-End Speech Synthesis
#8This seems like it could be great for automatically generating audio books. Personally I would one day like to have a program that can read arbitrary text to me in a more or less human way, that would allow me to read papers for work while driving.
Re: Predicting Expressive Speaking Style from Text in End-To-End Speech Synthesis
#9This seems like it could be great for automatically generating audio books. Personally I would one day like to have a program that can read arbitrary text to me in a more or less human way, that would allow me to read papers for work while driving.
I thought the exact same thing. Except that I know if it pronounces certain names or words wrong over and over it would make me crazy and I would have to stop.
Most decent TTS or assistive technology systems have a pronunciation dictionary. This is a very real problem for people who use screen readers on a daily basis but luckily it's a (mostly) solved one.
Re: Predicting Expressive Speaking Style from Text in End-To-End Speech Synthesis
#10This seems like it could be great for automatically generating audio books. Personally I would one day like to have a program that can read arbitrary text to me in a more or less human way, that would allow me to read papers for work while driving.
You could except you would get sued. The Kindle 2 was announced with a feature that would read the book to you and Amazon landed in court. https://sunsteinlaw.com/read-it-aloud-and-weep-controversy-s...
But yet here we are, 9 years later, and the Kindle apps on Android, Windows and iOS support screen reader access to books. Those screen readers can use an array of voices, undoubtedly including the speech engines used in the original TTS feature written about here.