I have a machine with two GPUs and a frozen OS with a tensorflow python app that clones voices, and I'd say the quality passes if you run it through a phone bandpass filter.
I've had an idea to use propellerhead recycle to chop the output cloned voice into syllables, and then "play" the chopped parts in rhythm, through autotune.
The issue is you get Eifel 65 sounding autotune if your base vocals are monotonic or way off key. The only way I can think of fixing this is to use something like audacity's pitch changer that doesn't affect the speed of the sample - rough the lyrical tones in with audacity/recycle, then autotune it where it needs to go.
I'd like to say I'm too busy to get this workflow going, but mostly I'm lazy and someone else will do it first - and better - I can't improve the AI cloning software.