Hey, do yourself a favor and listen to the fun example: > [S1] Oh fire! Oh my goodness! What's the procedure? What to we do people? The smoke could be coming through an air duct! Seriously impressive. Wish I could direct link the audio. Kudos to the Dia team.
Show HN: Dia, an open-weights TTS model for generating realistic dialogue
81–90 of 202 posts
Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue
#82Is this Apache licensed or a custom one? The README contains this: > This project is licensed under the Apache License 2.0 - see the LICENSE file for details. > This project offers a high-fidelity speech generation model *intended solely for research and educational use*. The following uses are strictly forbidden: > Identity Misuse: Do not produce audio resembling real individuals without permission. > ... Specifical…
Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue
#83Sounds really good & human! Got a fair bit of unexpected artifacts though. e.g. 3 seconds hissing noise before dialogue. And music in background when I added (happy) in an attempt to control tone. Also don't understand how to control the S1 and S2 speakers...is it just random based on temp? > TODO Docker support Got this adapted pretty easily. Just latest nvidia cuda container, throw python and modules on it and chan…
The outputs are a bit unstable, might need to add cleaner training data and run longer training sessions. Hopefully we can do something like OAI Whisper and update with better performing checkpoints!
Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue
#84Earlier quoted context omitted.
For anyone who wants to listen, it's on this page: https://yummy-fir-7a4.notion.site/dia
A little overacted, it reminds me of the voice acting in those flash cartoons you'd see in the early days of YouTube. That's not to say it isn't good work, it still sounds remarkably human. Just silly humans :)
Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue
#85Does this use the the mimi codec by moshi? If so it would be straighforward to get Dia running on iOS!
Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue
#86Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue
#87Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue
#88Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue
#89Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue
#90Earlier quoted context omitted.
A little overacted, it reminds me of the voice acting in those flash cartoons you'd see in the early days of YouTube. That's not to say it isn't good work, it still sounds remarkably human. Just silly humans :)
Reminded me of the Fenslerfilm G.I. Joe sketch where the kids have something on the stove burning