Live data from Hacker News

Show HN: Dia, an open-weights TTS model for generating realistic dialogue

github.com

21–30 of 202 posts

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#23
post #2

Hey HN! We’re Toby and Jay, creators of Dia. Dia is 1.6B parameter open-weights model that generates dialogue directly from a transcript. Unlike TTS models that generate each speaker turn and stitch them together, Dia generates the entire conversation in a single pass. This makes it faster, more natural, and easier to use for dialogue generation. It also supports audio prompts — you can condition the output on a spec…

Easily 10 times better than recent OpenAI voice model. I don't like robotic voices.

Example voices seems like over loud, over excitement like Andrew Tate, Speed or advertisement. It's lacking calm, normal conversation or normal podcast like interaction.

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#24
post #2

Hey HN! We’re Toby and Jay, creators of Dia. Dia is 1.6B parameter open-weights model that generates dialogue directly from a transcript. Unlike TTS models that generate each speaker turn and stitch them together, Dia generates the entire conversation in a single pass. This makes it faster, more natural, and easier to use for dialogue generation. It also supports audio prompts — you can condition the output on a spec…

Amazing that you developed this over the course of three months! Can you drop any insight into how you pulled together the audio data?

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#25
post #5
post #3

just in case, another opensource project using same name https://wiki.gnome.org/Apps/Dia/ https://gitlab.gnome.org/GNOME/dia

Thanks for the heads-up! We weren’t aware of the GNOME Dia project. Since we focus on speech AI, we’ll make sure to clarify that distinction.

Ditto this! Dia diagram tool user here just noticing the name clash. Good luck with your Dia!! Assuming both can exist in harmony. :-)

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#26
This is really impressive; we're getting close to a dream of mine: the ability to generate proper audiobooks from EPUBs. Not just a robotic single voice for everything, but different, consistent voices for each protagonist, with the LLM analyzing the text to guess which voice to use and add an appropriate tone, much like a voice actor would do.

I've tried "EPUB to audiobook" tools, but they are really miles behind what a real narrator accomplishes and make the audiobook impossible to engage with

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#27
post #3

just in case, another opensource project using same name https://wiki.gnome.org/Apps/Dia/ https://gitlab.gnome.org/GNOME/dia

I know it's a bit ridiculous to see that as some kind of conspiracy, but I have seen a very long list of AI-related projects that got the same name as a famous open-source project, as if they wanted to hijack the popularity of those projects, and Dia is yet another example. It was relatively famous a few years ago and you cannot have forgotten it if you used Linux for more than a few weeks. It's almost done on purpose.

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#28
Hey, do yourself a favor and listen to the fun example:

> [S1] Oh fire! Oh my goodness! What's the procedure? What to we do people? The smoke could be coming through an air duct!

Seriously impressive. Wish I could direct link the audio.

Kudos to the Dia team.

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#30
post #24
post #2

Hey HN! We’re Toby and Jay, creators of Dia. Dia is 1.6B parameter open-weights model that generates dialogue directly from a transcript. Unlike TTS models that generate each speaker turn and stitch them together, Dia generates the entire conversation in a single pass. This makes it faster, more natural, and easier to use for dialogue generation. It also supports audio prompts — you can condition the output on a spec…

Amazing that you developed this over the course of three months! Can you drop any insight into how you pulled together the audio data?

+1 to this, amazing how you managed to deliver this, and iff you're willing to share i'd be most interested in learning what you did in terms of train data..!
Post reply on HN