Live data from Hacker News

Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost

narilabs.com

21–30 of 37 posts

Re: Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost

#21

For some reason it switched voices half way through a 33 second clip. For OP the clip name is nari-nina-01a0a12f-980a-765e-8029-fa56bd23210d.wav

hey, thanks for letting us know! will look into the issue and see what went wrong.

Re: Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost

#22
post #12

Earlier quoted context omitted.

Yes, and it is very good one. Leading position on private leaderboard on HF: https://huggingface.co/spaces/hf-audio/open_asr_leaderboard

I meant the Nari inference engine for Qwen3-ASR. I'm aware that Qwen3-ASR is open source, but I don't see a repo under https://github.com/nari-labs for nari-qwen3-asr or similar. The Huggingface link on https://narilabs.com/product/stt/ links to https://huggingface.co/Qwen/Qwen3-ASR-1.7B , not anything under https://huggingface.co/nari-labs

the qwen3-asr inference repo is not OSSed as of now. we're planning to write a paper or tech report on it as it contains some general techniques for ASR inference.

Re: Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost

#23
post #22

Earlier quoted context omitted.

I meant the Nari inference engine for Qwen3-ASR. I'm aware that Qwen3-ASR is open source, but I don't see a repo under https://github.com/nari-labs for nari-qwen3-asr or similar. The Huggingface link on https://narilabs.com/product/stt/ links to https://huggingface.co/Qwen/Qwen3-ASR-1.7B , not anything under https://huggingface.co/nari-labs

the qwen3-asr inference repo is not OSSed as of now. we're planning to write a paper or tech report on it as it contains some general techniques for ASR inference.

how do I follow you? I have a small 5090 doing inference all the time and I barely use tts but a lot of asr, mostly whisper, I ported your tech report for tts and implemented some improvements on my whisper inference based on your tech report as well!

would love to talk sometime!

Re: Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost

#25
post #24

https://apimade.com/audio-compare.html Added it to my blind TTS model comparison leaderboard. So far Darwin TTS is the open model leading the pack, ElevenLabs is at the lead.

Is Darwin TTS from Fish Audio? It wasn't clear when I searched for it.

Re: Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost

#26
post #24

https://apimade.com/audio-compare.html Added it to my blind TTS model comparison leaderboard. So far Darwin TTS is the open model leading the pack, ElevenLabs is at the lead.

Is Darwin TTS from Fish Audio? It wasn't clear when I searched for it.

Darwin TTS is based on Qwen3-TTS

https://huggingface.co/zeropointnine/Darwin-TTS-1.7B-Cross-Q...

Re: Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost

#27
This is awesome. Thanks for pushing the audio pareto frontier forward.

Probably far fetched for now, but I think the next big evolution is building the pareto/much cheaper alternative to GPT-Live-1.

The STT/TTS market is quite saturated, while today, there's almost no cheap/open source alternative to GPT-Live-1.

Re: Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost

#28
post #24

https://apimade.com/audio-compare.html Added it to my blind TTS model comparison leaderboard. So far Darwin TTS is the open model leading the pack, ElevenLabs is at the lead.

awesome! will look into Darwin TTS. super interesting

Re: Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost

#29
post #27

This is awesome. Thanks for pushing the audio pareto frontier forward. Probably far fetched for now, but I think the next big evolution is building the pareto/much cheaper alternative to GPT-Live-1. The STT/TTS market is quite saturated, while today, there's almost no cheap/open source alternative to GPT-Live-1.

agreed. we've been doing some work around NVIDIA personaplex 7b, but its quality is quite far from GPT-Live-1, esp in terms of intelligence. Once a good OSS model is out, we'll be sure to be the first to serve it cheaply to the masses :)
Post reply on HN