I was wondering, wouldn't it be possible to classify the voice of every celebrity based on moods so one could make the voices less monotonic? So one could then add text metadata for the text-to-speech conversion, e.g. "[Angry] I have a dream, [Calm] but it has a patent so you can't copy it! (laughter) [Calm-fade-to-angry] In reality insomnia took it from me!"
The problem is that currently your training data has to be annotated with these tokens, and that adds a lot to the difficulty of creating data sets.
I imagine that over time this will get much easier to do.