Live data from Hacker News

Show HN: Sparrow-1 – Audio-native model for human-level turn-taking without ASR

tavus.io

11–20 of 50 posts

Re: Show HN: Sparrow-1 – Audio-native model for human-level turn-taking without ASR

#11
post #8

I am always skeptical of benchmarks that show perfect scores, especially when they come from the company selling the product. It feels like everyone claims to have solved conversational timing these days. I guess we will see if it is actually any good.

Different industry, but our marketing guy once said "You know what this [perfect] metric means? We can never use it in marketing because it's not believable"

Just include some noise, it’s like the most available resource in the universe

Re: Show HN: Sparrow-1 – Audio-native model for human-level turn-taking without ASR

#14

Literally no way to sign up to try. Put my email and password and it puts me into some wait list despite the video saying I could try the model today. That's what makes me mad about these kind of releases is that the marketing and the product don't talk together.

try signing up for the API platform on the site. You can access it there

Re: Show HN: Sparrow-1 – Audio-native model for human-level turn-taking without ASR

#15

Metric | Sparrow-1 Precision 100% Recall 100% Common ...

If you watch the demo video you can see how they would get this: the model is not aggressive enough. While it doesn't cut you off, which is nice, it also always waits an uncanny amount of time to chime in.

Re: Show HN: Sparrow-1 – Audio-native model for human-level turn-taking without ASR

#16

Metric | Sparrow-1 Precision 100% Recall 100% Common ...

If you watch the demo video you can see how they would get this: the model is not aggressive enough. While it doesn't cut you off, which is nice, it also always waits an uncanny amount of time to chime in.

That should lead to a low recall: too many false negatives. I wonder how they are calculating it.

Re: Show HN: Sparrow-1 – Audio-native model for human-level turn-taking without ASR

#17
post #9

I tried talking to Claude today. What a nightmare. It constantly interrupts you. I don’t mind if Claude wants to spend ten seconds thinking about its reply, but at least let ME finish my thought. Without decent turn-taking, the AI seems impolite and it’s just an icky experience. I hope tech like this gets widely distributed soon because there are so many situations in which I would love to talk with a model. If only…

Anthropic doesn't have any realtime multimodal audio models available, they just use STT and TTS models slapped on top of Claude. So they are currently the worst provider if you actually want to use voice communication.

Re: Show HN: Sparrow-1 – Audio-native model for human-level turn-taking without ASR

#19
post #11
post #8

Earlier quoted context omitted.

Different industry, but our marketing guy once said "You know what this [perfect] metric means? We can never use it in marketing because it's not believable"

Just include some noise, it’s like the most available resource in the universe

Never thought of noise as a resource, but yea.
Post reply on HN