This looks extremely similar to what often happened with chat implementations in 3 and 3.5 era: GPT world generate it's answer and then go on and generate the next input by the user as well.
The user's voice is vectorized somehow, and the model is predicting a series of vectors and won't have a sense of "self" to let it recognize it's own voice vs the user's.