The baseline configurations all note Is that really where SOTA is right now?
Absolutely not. 500-1000ms is borderline acceptable. Sub-300ms is closer to SOTA. 2000ms or more means people will hang up.
Asterisk AI Voice Agent
61–70 of 129 posts
Re: Asterisk AI Voice Agent
#62Earlier quoted context omitted.
Good customer support lines? Is there a reason why it can't provide good support. I often use chatgpt's voice function.
How? Businesses will use this to justify removing what few actual human support staff they have left. Nobody, and I mean it, nobody calls customer support because they want to talk to a computer. It’s the last resort for problems that usually can’t be accomplished via the already existing technical flows available via computer.
I wonder what Amazon's goals are, as an example. Currently, at least on the .ca website, there is no way to even get to chat to fix problems. All their spider text of help options, now always lead back to the return page.
So it's call them (which you can only find the number via Google.)
I suspect they're so disfunctional, that they don't understand why the massive uptick in calls, so then they slap AI in via phone too.
And so now that's slow and AI drivel. I guess soon I'll just have to do chargebacks!? Eg, if a package is missing or whatever.
Re: Asterisk AI Voice Agent
#63What is the application of this that makes anything better for anyone? All I can think of is more spammers, scammers, horrible customer support lines.
Re: Asterisk AI Voice Agent
#64Earlier quoted context omitted.
Good customer support lines? Is there a reason why it can't provide good support. I often use chatgpt's voice function.
How? Businesses will use this to justify removing what few actual human support staff they have left. Nobody, and I mean it, nobody calls customer support because they want to talk to a computer. It’s the last resort for problems that usually can’t be accomplished via the already existing technical flows available via computer.
Re: Asterisk AI Voice Agent
#65Earlier quoted context omitted.
How? Businesses will use this to justify removing what few actual human support staff they have left. Nobody, and I mean it, nobody calls customer support because they want to talk to a computer. It’s the last resort for problems that usually can’t be accomplished via the already existing technical flows available via computer.
That's not true. I recently called to make an appointment. I don't care if it's an AI. I would actually prefer it, because I wouldn't feel bad about taking a long time to pick the best time. Don't you think you're being a bit dogmatic about this?
Re: Asterisk AI Voice Agent
#66What is the application of this that makes anything better for anyone? All I can think of is more spammers, scammers, horrible customer support lines.
For narrow use cases like this I personally don't mind these tools.
Re: Asterisk AI Voice Agent
#67Earlier quoted context omitted.
Absolutely not. 500-1000ms is borderline acceptable. Sub-300ms is closer to SOTA. 2000ms or more means people will hang up.
play "Just a second, one moment please ".wave as soon as input goes quiet. ChatGPT app has a audio version of the spinner icon when you ask it a question and it needs a second before answering.
play "ehh".wavRe: Asterisk AI Voice Agent
#68Earlier quoted context omitted.
One easy way to build voice agents and connect them to Twilio is the Pipecat open source framework. Pipecat supports a wide variety of network transports, including the Twilio MediaStream WebSocket protocol so you don't have to bounce through a SIP server. Here's a getting started doc.[1] (If you do need SIP, this Asterisk project looks really great.) Pipecat has 90 or so integrations with all the models/services peo…
This is good stuff. In your opinion, how close is Pipecat + OSS to replacing proprietary infra from Vapi, Retell, Sierra, etc?
The integrated developer experience is much better on Vapi, etc.
The goal of the Pipecat project is to provide state of the art building blocks if you want to control every part of the multimodal, realtime agent processing flow and tech stack. There are thousands of companies with Pipecat voice agents deployed at scale in production, including some of the world's largest e-commerce, financial services, and healthtech companies. The Smart Turn model benchmarks better than any of the proprietary turn detection models. Companies like Modal have great info about how to build agents with sub-second voice-to-voice latency.[1] Most of the next-generation video avatar companies are building on Pipecat.[2] NVIDIA built the ACE Controller robot operating system on Pipecat.[3]
[1] https://modal.com/blog/low-latency-voice-bot - [2] https://lemonslice.com/ = [3] https://github.com/NVIDIA/ace-controller/
Re: Asterisk AI Voice Agent
#69Even if the focus is now on hosted telephony, my experience is that everywhere you can hear the default nusic-on-hold
Re: Asterisk AI Voice Agent
#70That seems like bad news for Allison. Though I know she already had some TTS voices available, so many not.