> Non-verbal cues are invisible to text: Transcription-based models discard sighs, throat-clearing, hesitation sounds, and other non-verbal vocalizations that carry critical conversational-flow information. Sparrow-1 hears what ASR ignores. Could Sparrow instead be used to produce high quality transcription that incorporate non-verbal cues? Or even, use Sparrow AND another existing transcription/ASR thing to augment…
Show HN: Sparrow-1 – Audio-native model for human-level turn-taking without ASR
31–40 of 50 posts
Re: Show HN: Sparrow-1 – Audio-native model for human-level turn-taking without ASR
#32Any examples available? Sounds amazing.
Re: Show HN: Sparrow-1 – Audio-native model for human-level turn-taking without ASR
#33How do I try the demo for Sparrow-1? What is pricing like?
Re: Show HN: Sparrow-1 – Audio-native model for human-level turn-taking without ASR
#34Metric | Sparrow-1 Precision 100% Recall 100% Common ...
The turn taking models were evaluated in a controlled environment with no additional cascaded steps: LLM, TTS, Phx. This matters to get apples to apples comparison: without the rest of the pipeline variability influencing the measurements.
The video conversation examples are sparrow-1 within the full pipeline. These responses aren’t as fast as sparrow itself because the LLM, TTS, facial rendering, and network transport also take time. Without Sparrow-1 they would be slower. Sparrow-1 enables the responses being as fast as they are, and with a faster CVI pipeline configuration the responses can be as fast as 430ms in my testing.
Re: Show HN: Sparrow-1 – Audio-native model for human-level turn-taking without ASR
#35Such things were doing a good-enough job scamming the elderly as it is--even with the silence-based delays.
Re: Show HN: Sparrow-1 – Audio-native model for human-level turn-taking without ASR
#36I tried talking to Claude today. What a nightmare. It constantly interrupts you. I don’t mind if Claude wants to spend ten seconds thinking about its reply, but at least let ME finish my thought. Without decent turn-taking, the AI seems impolite and it’s just an icky experience. I hope tech like this gets widely distributed soon because there are so many situations in which I would love to talk with a model. If only…
Re: Show HN: Sparrow-1 – Audio-native model for human-level turn-taking without ASR
#37I tried talking to Claude today. What a nightmare. It constantly interrupts you. I don’t mind if Claude wants to spend ten seconds thinking about its reply, but at least let ME finish my thought. Without decent turn-taking, the AI seems impolite and it’s just an icky experience. I hope tech like this gets widely distributed soon because there are so many situations in which I would love to talk with a model. If only…
Am I not allowed to cut you off if you're ramble-y and incoherent?
Re: Show HN: Sparrow-1 – Audio-native model for human-level turn-taking without ASR
#38Re: Show HN: Sparrow-1 – Audio-native model for human-level turn-taking without ASR
#39I tried talking to Claude today. What a nightmare. It constantly interrupts you. I don’t mind if Claude wants to spend ten seconds thinking about its reply, but at least let ME finish my thought. Without decent turn-taking, the AI seems impolite and it’s just an icky experience. I hope tech like this gets widely distributed soon because there are so many situations in which I would love to talk with a model. If only…
===
ME: "OK, so, I have a question about the economics of medicine. Uh..." [pauses to gather thoughts to ask question]
GEMINI: "Sure! Medical economics is the field of..."
===
And it's aggravated by the fact that all the LLMs love to give you page-long responses before it's your turn to talk again!
Re: Show HN: Sparrow-1 – Audio-native model for human-level turn-taking without ASR
#40I tried talking to Claude today. What a nightmare. It constantly interrupts you. I don’t mind if Claude wants to spend ten seconds thinking about its reply, but at least let ME finish my thought. Without decent turn-taking, the AI seems impolite and it’s just an icky experience. I hope tech like this gets widely distributed soon because there are so many situations in which I would love to talk with a model. If only…
I love Anthropic's models but their realtime voice is absolutely terrible. Every time I use it there is at least once that I curse at it for interrupting me. My main use case for OpenAI/ChatGPT at this point is realtime voice chats. OpenAI has done a pretty great job w/ realtime (their realtime API is pretty fantastic out of the box... not perfect, but pretty fantastic and dead simple setup). I can have what feels li…