WhisperSpeech – An open source text-to-speech system built by inverting Whisper
1–10 of 119 posts
Re: WhisperSpeech – An open source text-to-speech system built by inverting Whisper
#2Re: WhisperSpeech – An open source text-to-speech system built by inverting Whisper
#3Re: WhisperSpeech – An open source text-to-speech system built by inverting Whisper
#4Re: WhisperSpeech – An open source text-to-speech system built by inverting Whisper
#5I was looking at video on training a custom voice with Piper, following a tutorial at https://www.youtube.com/watch?v=b_we_jma220 , and noticed how the datasets required metadata of the text for the source audio files. This training method by Collabora seems to automate that process and only requires an audio file for training.
Re: WhisperSpeech – An open source text-to-speech system built by inverting Whisper
#6[0] https://github.com/netease-youdao/EmotiVoice
[1] https://github.com/siraben/emotivoice-cli
[2] https://github.com/netease-youdao/EmotiVoice/wiki/Voice-Clon...
Re: WhisperSpeech – An open source text-to-speech system built by inverting Whisper
#7Interested to see how it performs for Mandarin Chinese speech synthesis, especially with prosody and emotion. The highest quality open source model I've seen so far is EmotiVoice[0], which I've made a CLI wrapper around to generate audio for flashcards.[1] For EmotiVoice, you can apparently also clone your own voice with a GPU, but I have not tested this.[2] [0] https://github.com/netease-youdao/EmotiVoice [1] https:…
Re: WhisperSpeech – An open source text-to-speech system built by inverting Whisper
#8Interested to see how it performs for Mandarin Chinese speech synthesis, especially with prosody and emotion. The highest quality open source model I've seen so far is EmotiVoice[0], which I've made a CLI wrapper around to generate audio for flashcards.[1] For EmotiVoice, you can apparently also clone your own voice with a GPU, but I have not tested this.[2] [0] https://github.com/netease-youdao/EmotiVoice [1] https:…
Have you released your flashcard app?
Re: WhisperSpeech – An open source text-to-speech system built by inverting Whisper
#9Interested to see how it performs for Mandarin Chinese speech synthesis, especially with prosody and emotion. The highest quality open source model I've seen so far is EmotiVoice[0], which I've made a CLI wrapper around to generate audio for flashcards.[1] For EmotiVoice, you can apparently also clone your own voice with a GPU, but I have not tested this.[2] [0] https://github.com/netease-youdao/EmotiVoice [1] https:…
Re: WhisperSpeech – An open source text-to-speech system built by inverting Whisper
#10Am I weird in just having my head spin - even though I've also been at leading edge tech before, but this is just me yelling at these new algos on my lawn?