This is awesome. Thanks for pushing the audio pareto frontier forward. Probably far fetched for now, but I think the next big evolution is building the pareto/much cheaper alternative to GPT-Live-1. The STT/TTS market is quite saturated, while today, there's almost no cheap/open source alternative to GPT-Live-1.
Getting something conversationally better has been done, the tool calling will likely be worse though.
The infrastructure for real time is really annoying though.