Live data from Hacker News

How OpenAI delivers low-latency voice AI at scale

openai.com

131–140 of 172 posts

Re: How OpenAI delivers low-latency voice AI at scale

#135

The low latency is more of a pain point than a good thing, the way they have it implemented. Trying to have a casual conversation with it, as humans we naturally pause, and GPT will take this as you are "done" and start blabbing away. I also suffer from finding the appropriate word I want as I've gotten older and slower, and this fast-voice-gpt just ends up frustrating me more than helping. I have to sit there and th…

There's a really interesting project in Japanese natural language processing called J-Moshi that had a novel approach and in my opinion good results.

They tried to make it mimic the way Japanese is full of really quick acknowledgement sounds and it seems to allow it to handle those pauses and interruptions really well.

https://en.nagoya-u.ac.jp/news/articles/say-hello-to-j-moshi... (english)

https://nu-dialogue.github.io/j-moshi/ (japanese and english)

I must admit it's a bit weird when LLMs laugh, I don't really know how I feel about that but it seems to laugh at the right times. Very tangential, but cockatoos have been known to mimic the right time to laugh presumably based on tonal cues that a joke was just made (I have experienced this first hand with rescue birds who li e amongst humans)

Re: How OpenAI delivers low-latency voice AI at scale

#136

Wait a minute... I’m genuinely happy that they are sharing this, but keep in mind that realtime audio model from OpenAI are still stuck with the 4o family in terms of capabilities, sadly. I still find them so useful, such a pity that there’s no real competitor in this segment, having the experience a real conversation has helped me so much in expressing ideas and concepts. Still, it’s worth to keep in mind that these…

You can feel what is possible using Gemini speech to speech model, it can do tool calls and is very fast. It lacks somewhat in thinking capability but you can setup a tool call to a smarter model and it acts as a relay. I’ve been very impressed.

Re: How OpenAI delivers low-latency voice AI at scale

#137

Earlier quoted context omitted.

I’ve tried this and it says it will but just keeps cutting in. I hate this feature so much.

If anyone has an alternative I’m all ears. This would be a killer feature for me and something I’ve tried to use on cross-country road trips.

I know it's not the perfect solution for you, but I use a voice recorder and send the LLM the transcript. And my god is it working great.

Usually I just explain the things I want it to do. The longest was 30 minutes rambling of explaining the methods section of a paper in non chronological order. It worked unbelievable good for me.

Re: How OpenAI delivers low-latency voice AI at scale

#140

The low latency is more of a pain point than a good thing, the way they have it implemented. Trying to have a casual conversation with it, as humans we naturally pause, and GPT will take this as you are "done" and start blabbing away. I also suffer from finding the appropriate word I want as I've gotten older and slower, and this fast-voice-gpt just ends up frustrating me more than helping. I have to sit there and th…

[dead]
Post reply on HN