Live data from Hacker News

Show HN: Dia, an open-weights TTS model for generating realistic dialogue

github.com

171–180 of 202 posts

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#172

I inserted the non-verbal command "(pauses)" in the middle of a sentence and I think I caused it to have an aneurysm. https://i.horizon.pics/4sEVXh8GpI (27s) It starts with an intro, too. Really strange

That was... amazing.

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#173

Earlier quoted context omitted.

Why a human? There are many cases where I like a book but dislike the audiobook speaker, so I essentially can't listen to that book anymore. With a machine, I can tweak the voice to my heart's content.

And get a completely wrong/bland but custom read of the book. Reading is much more than simply transforming text to audio.

Sometimes, I don't care if it's bland, I just want to listen to the text. There are a lot of Asian light novels for example which never get English audiobooks, and I've listened to many of them with basic TTS, not even an AI model TTS like these more recent ones, and I thoroughly enjoyed these books even still.

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#174
post #112

Anyone know if possible to fine-tune for cloning my voice?

We're adding guides for Zero-shot voice cloning. You can try it using the second example on Gradio: https://huggingface.co/spaces/nari-labs/Dia-1.6B

Will give it a shot but I feel like fine-tuning will be more reliable, any way to do that?

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#175
Seeing is no longer believing. Hearing isn't either. The funny thing is, it's getting to the point where LLM-generated text is more easily spotted than AI audio, video, and images.

It's going to be an interesting decade of the new equivalent of "No, Tiffany, Bill Gates will NOT be sending you $100 for forwarding that email." Except it's going to be AI celebrities making appeals for donations to help them become billionaires or something.

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#177

Hey, do yourself a favor and listen to the fun example: > [S1] Oh fire! Oh my goodness! What's the procedure? What to we do people? The smoke could be coming through an air duct! Seriously impressive. Wish I could direct link the audio. Kudos to the Dia team.

Yeah, that example is insane.

Is there some sort of system prompt or hint at how it should be voiced, or does it interpret it from the text?

Because it would be hilarious if it just derived it from the text and it did this sort of voice acting when you didn't want it to, like reading a matter-of-fact warning label.

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#178
post #48

Hey, do yourself a favor and listen to the fun example: > [S1] Oh fire! Oh my goodness! What's the procedure? What to we do people? The smoke could be coming through an air duct! Seriously impressive. Wish I could direct link the audio. Kudos to the Dia team.

For anyone who wants to listen, it's on this page: https://yummy-fir-7a4.notion.site/dia

Sounds great. One of the female examples has convincing uptalk. There must be a way to manipulate the latent space to control uptalk, vocal fry, smoker’s voice, lispiness, etc.

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#179

I inserted the non-verbal command "(pauses)" in the middle of a sentence and I think I caused it to have an aneurysm. https://i.horizon.pics/4sEVXh8GpI (27s) It starts with an intro, too. Really strange

I have a hunch they're pulling data from radio shows to give it that "high quality" vibe. Tried running it through this script and hit some weird bugs too: [S1] It really sounds as if they've started using NPR to source TTS models [S2] Yeah... yeah... it's kind of disturbing (laughs dejectedly). [S3] I really wish, that they would just Stop with this. https://i.horizon.pics/Tx2PrPTRM3

It even added an extra f-word at the end. Still veeery impressive

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#180
post #51
post #31

Earlier quoted context omitted.

The generous interpretation is that the AI hype people just didn’t know about those other projects, i.e. that they are neither open source developers, nor users.

Of course, how could they have known? Doing a basic web search before deciding on a name is so last year.

Maybe they only asked an LLM about it?
Post reply on HN