I think we're close to the point where local models can make this economically viable as an alternative to hands off XAI/OAI/ANT
You need rapid good enough rapid turn response + rapid good enough TTS in order to make it economically for someone who's not blending it into a more compute heavy pricing model
Have you looked into it? Are you familiar with Pipecat https://github.com/pipecat-ai/pipecat - they put out interesting demos frequently