"Once your request is submitted, it takes one to two months to process the book and conduct quality checks." My guess is that these generated voices are far from perfect and someone has to go in and crank the algorithm to get a fair number of passages to not sound strange. Even in the example Helena there is a word at the end of a sentence that sounds like it should be in the middle and has a bit of weirdness to it.…
Why is that we still can't have a perfect or near-perfect text-to-speech given all the astonishing advances in ML taking place? Is TTS an area nobody is really interested in or is it harder than generating beautiful pictures and sophisticated writings? This thing by Apple already sounds way better than the best I heard previously (NextUp Ivona) but it is not an instant-result offline tool yet and that's sad.
I think one differences with pictures and audio is that pictures are two-dimensional and we can't take in the whole image at a time. This makes it easy to overlook flaws without careful inspection. And I find that although there has been some amazing AI-generated art, there are still a lot of rough edges and tweaking required to get really clean images.
As far as writing goes, I suspect that the rules of written language are easier to learn and violations easier to overlook than with generated audio.