Hey, do yourself a favor and listen to the fun example: > [S1] Oh fire! Oh my goodness! What's the procedure? What to we do people? The smoke could be coming through an air duct! Seriously impressive. Wish I could direct link the audio. Kudos to the Dia team.
For anyone who wants to listen, it's on this page: https://yummy-fir-7a4.notion.site/dia
Show HN: Dia, an open-weights TTS model for generating realistic dialogue
61–70 of 202 posts
Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue
#62Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue
#63This is really impressive; we're getting close to a dream of mine: the ability to generate proper audiobooks from EPUBs. Not just a robotic single voice for everything, but different, consistent voices for each protagonist, with the LLM analyzing the text to guess which voice to use and add an appropriate tone, much like a voice actor would do. I've tried "EPUB to audiobook" tools, but they are really miles behind wh…
Wouldn’t it be more desirable to hear an actual human on an audiobook? Ideally the author?
The author could be better, because they at least have other info beyond the text to rely on, they can go off-script or add little details, etc.
Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue
#64Hey, do yourself a favor and listen to the fun example: > [S1] Oh fire! Oh my goodness! What's the procedure? What to we do people? The smoke could be coming through an air duct! Seriously impressive. Wish I could direct link the audio. Kudos to the Dia team.
This is so good. Reminds me of The Office. I love how bad the other examples are.
Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue
#65made a small change and got it running on M2 Pro 16GB Macbook pro, the quality is amazing. https://github.com/nari-labs/dia/pull/4
Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue
#66> TODO Docker support
Got this adapted pretty easily. Just latest nvidia cuda container, throw python and modules on it and change server to serve on 0.0.0.0. Does mean it pulls the model every time on startup though which isn't ideal
Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue
#67I've just been massively disappointed by Sesame's CSM: on their gradio on the website it was generating flawless dialogs with amazing voice cloning. When running it local the voice cloning performance is awful.
Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue
#68Earlier quoted context omitted.
Wouldn’t it be more desirable to hear an actual human on an audiobook? Ideally the author?
Honestly, I’d say that’s true only for the author. Anyone else is just going to be interpreting the words to understand how to best convey the character / emotion / situation / etc., just like an AI will have to do. If an AI can do that more effectively than a human, why not? The author could be better, because they at least have other info beyond the text to rely on, they can go off-script or add little details, etc…
The most skilled readers will make you want to read books _just because they narrated them_. They add a unique quality to the story, that you do not get from reading yourself or from watching a video adaptation.
Currently I'm in The Age of Madness, read by Steven Pacey. He's fantastic. The late Roy Dotrice is worth a mention as well, for voicing Game of Thrones and claiming the Guinness world record for most distinct voices (224) in one series.
It will be awesome if we can create readings automatically, but it will be a while before TTS can compete with the best readers out there.
Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue
#69Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue
#70This is really impressive; we're getting close to a dream of mine: the ability to generate proper audiobooks from EPUBs. Not just a robotic single voice for everything, but different, consistent voices for each protagonist, with the LLM analyzing the text to guess which voice to use and add an appropriate tone, much like a voice actor would do. I've tried "EPUB to audiobook" tools, but they are really miles behind wh…
Wouldn’t it be more desirable to hear an actual human on an audiobook? Ideally the author?
Of course, but it's not always available.
For example, I would love an audiobook for Stanisław Lem's "The Invincible," as I just finished its video game adaptation, yet it simply doesn't exist in my native language.
It's quite seldom that the author narrates the audiobooks I listen to, and sometimes the narrator does a horrible job, butchering the characters with exaggerated tones.