How OpenAI delivers low-latency voice AI at scale
111–120 of 172 posts
Re: How OpenAI delivers low-latency voice AI at scale
#112The low latency is more of a pain point than a good thing, the way they have it implemented. Trying to have a casual conversation with it, as humans we naturally pause, and GPT will take this as you are "done" and start blabbing away. I also suffer from finding the appropriate word I want as I've gotten older and slower, and this fast-voice-gpt just ends up frustrating me more than helping. I have to sit there and th…
yeh exactly, you cannot get a strong signal that a user is done speaking without some amount of “wait for 500ms of silence”. You could kick of processing and abandon if they continued talking, but that seems over optimized. 1-2s replies feel natural and like you pointed out pausing for 2-3s mid sentence is super normal.
Re: How OpenAI delivers low-latency voice AI at scale
#113The low latency is more of a pain point than a good thing, the way they have it implemented. Trying to have a casual conversation with it, as humans we naturally pause, and GPT will take this as you are "done" and start blabbing away. I also suffer from finding the appropriate word I want as I've gotten older and slower, and this fast-voice-gpt just ends up frustrating me more than helping. I have to sit there and th…
Re: How OpenAI delivers low-latency voice AI at scale
#114Earlier quoted context omitted.
Check out [0]. You can do 'Voice AI' on small/cheap hardware. It's the most fun you can have in the space ATM :) It's been a while, but posted a demo here [1] [0] https://github.com/pipecat-ai/pipecat-esp32 [1] https://www.youtube.com/watch?v=6f0sUEUuruw
beautiful demo - is it running fully locally or talking to 3rd party API’s? That box was jaw dropping small
Re: How OpenAI delivers low-latency voice AI at scale
#115Earlier quoted context omitted.
Is it reductive when its describing a group of people that like something and refusing to hear any ill of it? The comment wasn't shade at people using the language in general. And you're right, fanboys are in every language. But resorting to changing the argument by whataboutism is a bit reductive.
I’m not a go fanboy, but I do know from other contexts that so-called “fanboy“ behaviour is frequently associated with level-headed supporters getting defensive in the face of imprecise criticism. There’s an oft-repeated pattern where valid specific criticisms morph into broad criticism, which morphs into judgement, which breeds defensiveness, which feeds the criticism. Once you recognise this pattern, you see it eve…
Re: How OpenAI delivers low-latency voice AI at scale
#116The low latency is more of a pain point than a good thing, the way they have it implemented. Trying to have a casual conversation with it, as humans we naturally pause, and GPT will take this as you are "done" and start blabbing away. I also suffer from finding the appropriate word I want as I've gotten older and slower, and this fast-voice-gpt just ends up frustrating me more than helping. I have to sit there and th…
I’ve also experienced this and it’s really annoying. There is this pressure to keep talking if I’m not done with my thought that feels pretty unnatural at least for me. If I’m searching for the right word, I want the opportunity to find it. I think the solution is to handle pauses more intelligently rather than having a higher latency protocol. With low latency you can interrupt and the bot can immediately stop rambl…
Re: How OpenAI delivers low-latency voice AI at scale
#117Earlier quoted context omitted.
I’ve also experienced this and it’s really annoying. There is this pressure to keep talking if I’m not done with my thought that feels pretty unnatural at least for me. If I’m searching for the right word, I want the opportunity to find it. I think the solution is to handle pauses more intelligently rather than having a higher latency protocol. With low latency you can interrupt and the bot can immediately stop rambl…
100%. I have to hold the floor by filling the space with "ummmmmmmm.... uhhhh...." which inevitably distracts me from my point altogether. Poor user experience.
Re: How OpenAI delivers low-latency voice AI at scale
#118Re: How OpenAI delivers low-latency voice AI at scale
#119Earlier quoted context omitted.
Yeah, that's why they've used "reach" - the total number of users who could be exposed to the feature regardless of engagement.
To defend them a little: voice is a little rough around the edges now, so there’s a chicken and egg problem of whether to prioritize improving voice if usage isn’t high partially because it’s clunky.
Re: How OpenAI delivers low-latency voice AI at scale
#120Wait a minute... I’m genuinely happy that they are sharing this, but keep in mind that realtime audio model from OpenAI are still stuck with the 4o family in terms of capabilities, sadly. I still find them so useful, such a pity that there’s no real competitor in this segment, having the experience a real conversation has helped me so much in expressing ideas and concepts. Still, it’s worth to keep in mind that these…
Grok voice is surprisingly good, actually. It's still a dumber model than the thinking modes of frontier models, but it's less dumb than the voice modes of other providers.
Just give me a option to have a slower response but better model…