Live data from Hacker News

Show HN: Dia, an open-weights TTS model for generating realistic dialogue

github.com

161–170 of 202 posts

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#161
post #48

Earlier quoted context omitted.

For anyone who wants to listen, it's on this page: https://yummy-fir-7a4.notion.site/dia

A little overacted, it reminds me of the voice acting in those flash cartoons you'd see in the early days of YouTube. That's not to say it isn't good work, it still sounds remarkably human. Just silly humans :)

Overacted and silly humans indeed: https://www.youtube.com/watch?v=gO8N3L_aERg

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#162

I inserted the non-verbal command "(pauses)" in the middle of a sentence and I think I caused it to have an aneurysm. https://i.horizon.pics/4sEVXh8GpI (27s) It starts with an intro, too. Really strange

I have a hunch they're pulling data from radio shows to give it that "high quality" vibe. Tried running it through this script and hit some weird bugs too:

    [S1] It really sounds as if they've started using NPR to source TTS models
    [S2] Yeah... yeah... it's kind of disturbing (laughs dejectedly).
    [S3] I really wish, that they would just Stop with this.
https://i.horizon.pics/Tx2PrPTRM3

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#163

Very cool! Insane how much low hanging fruit there is for Audio models right now. A team of two picking things up over a few months can build something that still competes with large players with tons of funding

Yeah, Eleven Labs must be raking it in. You can get hours of audio out of it for free with Eleven Reader, which suggests that their inference costs aren't that high. Meanwhile, those same few hours of audio, at the exact same quality, would cost something like $100 when generated through their website or API, a lot more than any other provider out there. Their pricing (and especially API pricing) makes no sense, not…

Kokoro gives great results especially when speaking english. Model is small enough to run even on smartphone ~3x faster than realtime.

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#164
post #163

Earlier quoted context omitted.

Yeah, Eleven Labs must be raking it in. You can get hours of audio out of it for free with Eleven Reader, which suggests that their inference costs aren't that high. Meanwhile, those same few hours of audio, at the exact same quality, would cost something like $100 when generated through their website or API, a lot more than any other provider out there. Their pricing (and especially API pricing) makes no sense, not…

Kokoro gives great results especially when speaking english. Model is small enough to run even on smartphone ~3x faster than realtime.

Another +1 to Kokoro from me, great quality with good speed.

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#165

I inserted the non-verbal command "(pauses)" in the middle of a sentence and I think I caused it to have an aneurysm. https://i.horizon.pics/4sEVXh8GpI (27s) It starts with an intro, too. Really strange

That's certainly unusual ...

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#167
post #48

Hey, do yourself a favor and listen to the fun example: > [S1] Oh fire! Oh my goodness! What's the procedure? What to we do people? The smoke could be coming through an air duct! Seriously impressive. Wish I could direct link the audio. Kudos to the Dia team.

For anyone who wants to listen, it's on this page: https://yummy-fir-7a4.notion.site/dia

This is an instant classic. Sesame comparison examples all sound like clueless rich people from The White Lotus.

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#168
post #120

Hey, this is really cool! Curious how good the multi-language support is. Also - pretty wild that you trained the whole thing yourselves, especially without prior experience in speech models. Might actually be helpful for others if you ever feel like documenting how you got started and what the process looked like. I’ve never worked with TTS models myself, and honestly wouldn’t know where to begin. Either way, awesom…

Thank you so much for the kind words :) We only support English at the moment, hopefully can do more languages in the future. We are planning to release a technical report on some of the details, so stay tuned for that!

I'd also love to peek behind the curtains, if only to satisfy my own curiosity. Looking forward to the technical report, well done!

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#170

I inserted the non-verbal command "(pauses)" in the middle of a sentence and I think I caused it to have an aneurysm. https://i.horizon.pics/4sEVXh8GpI (27s) It starts with an intro, too. Really strange

You hittin them balloons again, mate ?
Post reply on HN