Yes, it assumes an attacker cannot imitate the other person's voice without detection. If you introduce enough latency or have a system fast enough to do on-the-fly voice changing you'd be set. It's such a great, simple idea, I felt stupid for not having thought of it.
To be sure, the voice channel establishes before the verbal authentication. I see your fingerprint is "banana", you see "kitchen". We can chat and I can say "alright so banana, right?" and you say "yep, in the kitchen". An attacker would need to remove banana and kitchen from that voice conversation and put their own fingerprint words in.
I think that level of realtime audio modification is pretty out of reach for now, although I suppose if you introduced a lot of latency, you might be able to pull it off. It'd probably be noticeable, and you can keep chatting and confirming the phrase, so an attacker would have to be really on-the-ball.
Plus, this only happens the first time. After that, the client saves the keys, so you know you're good.