this is cool but for agent scenarios unless an LLM bakes in the speech tokens directly, the latency is lost to inference, and this is what makes openai's voice model so interesting also sweet spot is under 150ms so the remainder is inference latency turn around, a 50ms turnaround including tts-stt would ofc be the dream that is "this ai agent is indistinguishably present and sentient" area
LLM TTFT is still a big issue, and we might tackle that problem as well.