Live data from Hacker News

Show HN: Dia, an open-weights TTS model for generating realistic dialogue

github.com

191–200 of 202 posts

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#191
post #163

Earlier quoted context omitted.

Yeah, Eleven Labs must be raking it in. You can get hours of audio out of it for free with Eleven Reader, which suggests that their inference costs aren't that high. Meanwhile, those same few hours of audio, at the exact same quality, would cost something like $100 when generated through their website or API, a lot more than any other provider out there. Their pricing (and especially API pricing) makes no sense, not…

Kokoro gives great results especially when speaking english. Model is small enough to run even on smartphone ~3x faster than realtime.

Kokoro just proves my point; it's "one guy in a garage", 1000 hours of distilled audio (I think) and ~100m params.

With the budget one tenth that of Stable Diffusion and less ethical qualms, you could easily 10x or 100x this.

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#192

I inserted the non-verbal command "(pauses)" in the middle of a sentence and I think I caused it to have an aneurysm. https://i.horizon.pics/4sEVXh8GpI (27s) It starts with an intro, too. Really strange

I have a hunch they're pulling data from radio shows to give it that "high quality" vibe. Tried running it through this script and hit some weird bugs too: [S1] It really sounds as if they've started using NPR to source TTS models [S2] Yeah... yeah... it's kind of disturbing (laughs dejectedly). [S3] I really wish, that they would just Stop with this. https://i.horizon.pics/Tx2PrPTRM3

The “Yeah…” followed by an uncomfortably long pause then a second “Yeah…” killed me.

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#193

This is really impressive; we're getting close to a dream of mine: the ability to generate proper audiobooks from EPUBs. Not just a robotic single voice for everything, but different, consistent voices for each protagonist, with the LLM analyzing the text to guess which voice to use and add an appropriate tone, much like a voice actor would do. I've tried "EPUB to audiobook" tools, but they are really miles behind wh…

Wouldn’t it be more desirable to hear an actual human on an audiobook? Ideally the author?

It'd be nice if there were mainstream releases on GBC/GBA/PSP again too! But apparently if there's no money in something then people don't really wanna do it.

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#194

Earlier quoted context omitted.

Why would that be a taboo question to ask? It should be the question we always ask, when presented with a model and in some cases we should probably reject the model, based on that information.

Because generally the person asking this question is trying to cancel the model maker

Well presumably since they're individuals and not a business the consequences are much less severe legally - but public opinion still won't be great, but since when was it ever, for any new thing?

If I cut up a song or TV show & put it on Youtube (and screech about fair use/parody law) then that's fine, but people will balk at something like this.

AI is here, people.

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#196

Earlier quoted context omitted.

I have a hunch they're pulling data from radio shows to give it that "high quality" vibe. Tried running it through this script and hit some weird bugs too: [S1] It really sounds as if they've started using NPR to source TTS models [S2] Yeah... yeah... it's kind of disturbing (laughs dejectedly). [S3] I really wish, that they would just Stop with this. https://i.horizon.pics/Tx2PrPTRM3

It even added an extra f-word at the end. Still veeery impressive

I think it also said "maaan" at the end of the previous line. And speaks out lout the "dejectedly".

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#197
post #140

Earlier quoted context omitted.

You're absolutely right. We used Jordan's Whisper-D, and he was generous enough to offer some guidance along the way. It's also a valid criticism that we haven’t yet audited the dataset for existing list of tags. That’s something we’ll be improving soon. As for Dia’s architecture, we largely followed existing models to build the 1.6B version. Since we only started learning about speech AI three months ago, we chose n…

I’m curious what differentiates it from Parakeet? I was listening to some of the demos on the parakeet announcement and they sound very similar to your examples - are they trained on the same data? Are there benefits to using Dia over Parakeet?

Well, this for one, about Parakeet:

> We plan to release our fine-tuned whisper models and possibly the generative model (and/or future improved versions). The generative model would have to be released under a non-commercial license due to our datasets.

https://jordandarefsky.com/blog/2024/parakeet/

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#198

Earlier quoted context omitted.

A little overacted, it reminds me of the voice acting in those flash cartoons you'd see in the early days of YouTube. That's not to say it isn't good work, it still sounds remarkably human. Just silly humans :)

"flash cartoons in the early days of Youtube" Wouldn't those be straight from Newgrounds?

Thank you! I couldn't remember the name Newgrounds for some reason!!

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#199
post #153

Earlier quoted context omitted.

Because generally the person asking this question is trying to cancel the model maker

No. It's for giving credit where credit is due. And yes, that includes the question if the people who generated the training data in the first place have given their consent that this can be used for AI training. It's quite concerning that the community around here is usually livid about FOSS license violations, which typically use copyright law as leverage, but somehow is perfectly OK with training models on copyrig…

What AI tools have you used recently? Have you verified if they all use models trained on copyrighted material with permission?

Re: Show HN: Dia, an open-weights TTS model for generating realistic dialogue

#200
post #153

Earlier quoted context omitted.

No. It's for giving credit where credit is due. And yes, that includes the question if the people who generated the training data in the first place have given their consent that this can be used for AI training. It's quite concerning that the community around here is usually livid about FOSS license violations, which typically use copyright law as leverage, but somehow is perfectly OK with training models on copyrig…

What AI tools have you used recently? Have you verified if they all use models trained on copyrighted material with permission?

Ah, that's a classic. "How can you criticize Big Oil and at the same time drive a car!" and voila, the case is closed.

I am allowed to criticize things without having to live like a hermit. I make moderate use of ChatGPT, yet at the same time I think that its training does not fall under fair use, and that creators should get compensated. If OpenAI's business model does not allow for this, then it should fail, and that's fine by me. I lived without ChatGPT, and I can live without it again.

Post reply on HN