Live data from Hacker News

Show HN: Three new Kitten TTS models – smallest less than 25MB

github.com

31–40 of 201 posts

Re: Show HN: Three new Kitten TTS models – smallest less than 25MB

#31

This would be great as a js package - 25mb is small enough that I think it'd be worth it (in-browser tts is still pretty bad and varies by browser)

great idea, we're on it. we're also working on a mobile sdk. a browser sdk would be really cool too.

Re: Show HN: Three new Kitten TTS models – smallest less than 25MB

#33

Earlier quoted context omitted.

thank you so much. Right now, it cannot handle expressive tags. what kind of tags would be most helpful according to you?

Emotion based tagging control would be the most helpful narrowing it down. Tags like [sarcastically] [happily] [joyfully] [fearfully]: so a subsection of adverbs. A stretch goal is 'arbitrary tags' from [singing] [sung to the tune of {x}] [pausing for emphasis] [slowly decreasing speed for emphasis] [emphasizing the object of this sentence] [clapping] [car crash in the distance] [laser's pew pew]. But yeah: instructi…

yeah i think to start with, narrowing it down to a few tags would be most helpful and we'll probably start w that first. Thanks a lot!

Re: Show HN: Three new Kitten TTS models – smallest less than 25MB

#34
post #23

A lot of these models struggle with small text strings, like "next button" that screen readers are going to speak a lot.

I think I tried on my Android everything I could try and 1. outside webpage reading, not many options; 2. as browser extensions, also not many (I don't like to copy URLs in your app) 3. they all insist reading every little shit, not only buttons but also "wave arrow pointing directly right" which some people use in their texts. So basically reading text aloud is a bunch of shitty options. Anyone jumping in this marke…

we'd love to serve this use-case. i'll make a demo for this next week and comment here with it.

Re: Show HN: Three new Kitten TTS models – smallest less than 25MB

#35

I'm still looking for the "perfect" setup in order to clone my voice and use it locally to send voice replies in telegram via openclaw. Does anyone have auch a setup? I want to be my own personal assistant... EDIT: I can provide it a RTX 3080ti.

Is it not just to train a model on your voice recordings and just use that to generate audio clips from text?

Re: Show HN: Three new Kitten TTS models – smallest less than 25MB

#38
post #3

Thanks for open sourcing this. Is there any way to do a custom voice as a DIY? Or we need to go through you? If so, would you consider making a pricing page for purchasing a license/alternative voice? All but one of the voices are unusable in a business context.

thanks a lot for the feedback. yes, we're working on a diy way to add custom voices and will also be releasing a model with more professional voices in the next 2-3 weeks. as of now, we're providing commercial support for custom voices, languages and deployment through the support form on our github. can you share more about your business use-case? if possible, i'd like to ensure the next release can serve that.

Right now it's outgoing calls for a small business client that checks information. Although if they call back they don't mind an automated system, on outgoing calls the person answering will often hang up if they detect AI right away, so we use a realistic custom voice with an accent.

This is a mind numbing task that requires workers to make hundreds of calls each day with only minor variations, sometimes navigating phone trees, half the time leaving almost the exact same message.

Anyway, I believe almost all such businesses will be automated within months. Human labour just cannot compete on cost.

Post reply on HN