Live data from Hacker News

Show HN: Dia, an open-weights TTS model for generating realistic dialogue

github.com

201–202 of 202 posts

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#201

I inserted the non-verbal command "(pauses)" in the middle of a sentence and I think I caused it to have an aneurysm. https://i.horizon.pics/4sEVXh8GpI (27s) It starts with an intro, too. Really strange

[pauses] i think i heard demons

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#202
post #163

Earlier quoted context omitted.

Kokoro gives great results especially when speaking english. Model is small enough to run even on smartphone ~3x faster than realtime.

Kokoro just proves my point; it's "one guy in a garage", 1000 hours of distilled audio (I think) and ~100m params. With the budget one tenth that of Stable Diffusion and less ethical qualms, you could easily 10x or 100x this.

I'm actually surprised people aren't just using elevenreader to generate solid content from various books for datasets lol
Post reply on HN