Is there any work/progress on a NN/trained model based on Internet Phonetic Alphabet (IPA) transcriptions? I.e. to be able to create an IPA transcription and convert it back to sound. That approach would be useful for things like shifting a voice to a different accent and to support voices speaking multiple languages. This can be done to a limited extend for models such as MBROLA voices by mapping the phonemes of one…
There is still a lot to explore in this space – we certainly don't have all the answers yet!