"Once your request is submitted, it takes one to two months to process the book and conduct quality checks." My guess is that these generated voices are far from perfect and someone has to go in and crank the algorithm to get a fair number of passages to not sound strange. Even in the example Helena there is a word at the end of a sentence that sounds like it should be in the middle and has a bit of weirdness to it.…
Why is that we still can't have a perfect or near-perfect text-to-speech given all the astonishing advances in ML taking place? Is TTS an area nobody is really interested in or is it harder than generating beautiful pictures and sophisticated writings? This thing by Apple already sounds way better than the best I heard previously (NextUp Ivona) but it is not an instant-result offline tool yet and that's sad.
Define perfect ;) Two different people will read the same text slightly (or not slightly) differently.
A great example is this brilliant and funny rendition of "To be or not to be" by Tim Minchin, Benedict Cumberbatch, Judy Dench, David Tennant and others. Sorry for the Facebook link, but it's very hard to find this video anywhere: https://www.facebook.com/watch/?v=585252039999241