One interesting application of tech like this is to produce story mods for games that still sound like they're using the original voice actors.
Is it possible to create a voice changer with these kind of AI?
It would basically involve a two-step approach where the first model extracts text and intonation and the second model synthesises the target voice.