Show HN: Three new Kitten TTS models – smallest less than 25MB
21–30 of 201 posts
Re: Show HN: Three new Kitten TTS models – smallest less than 25MB
#22One of the core features I look for is expressive control. Either in the form of the api via pitch/speed/volume controls, for more deterministic controls. Or in expressive tags such as [coughs], [urgently], or [laughs in melodic ascending and descending arpeggiated gibberish babbles]. the 25MB model is amazingly good for being 25MB. How does it handle expressive tags?
thank you so much. Right now, it cannot handle expressive tags. what kind of tags would be most helpful according to you?
A stretch goal is 'arbitrary tags' from [singing] [sung to the tune of {x}] [pausing for emphasis] [slowly decreasing speed for emphasis] [emphasizing the object of this sentence] [clapping] [car crash in the distance] [laser's pew pew].
But yeah: instruction/control via [tags] is the deciding feature for me, provided prompt adherence is strong enough.
Also: a thought...
Everyone is using [] for different kinds of tags in this space: which is very simple. Maybe it makes sense to differentiate kinds of tags? I.E. [tags for modifying how text is spoken] vs {tags for creating sounds not specifically speech: not modifying anything... but instead it's own 'sound/word'}
Re: Show HN: Three new Kitten TTS models – smallest less than 25MB
#23A lot of these models struggle with small text strings, like "next button" that screen readers are going to speak a lot.
Re: Show HN: Three new Kitten TTS models – smallest less than 25MB
#24A lot of good small TTS models in recent times. Most seem to struggle hard on prosody though. Kokoro TTS for example has a very good Norwegian voice but the rhythm and emphasizing is often so out of whack the generated speech is almost incomprehensible. Haven't had time to check this model out yet, how does it fare here? What's needed to improve the models in this area now that the voice part is more or less solved?
Re: Show HN: Three new Kitten TTS models – smallest less than 25MB
#25Re: Show HN: Three new Kitten TTS models – smallest less than 25MB
#26Re: Show HN: Three new Kitten TTS models – smallest less than 25MB
#27Re: Show HN: Three new Kitten TTS models – smallest less than 25MB
#28A lot of good small TTS models in recent times. Most seem to struggle hard on prosody though. Kokoro TTS for example has a very good Norwegian voice but the rhythm and emphasizing is often so out of whack the generated speech is almost incomprehensible. Haven't had time to check this model out yet, how does it fare here? What's needed to improve the models in this area now that the voice part is more or less solved?
Re: Show HN: Three new Kitten TTS models – smallest less than 25MB
#29[flagged]
The new 15M is way better than the previous 80M model(v0.1). So we're able to predictably improve the quality which is very encouraging.
Re: Show HN: Three new Kitten TTS models – smallest less than 25MB
#30A lot of good small TTS models in recent times. Most seem to struggle hard on prosody though. Kokoro TTS for example has a very good Norwegian voice but the rhythm and emphasizing is often so out of whack the generated speech is almost incomprehensible. Haven't had time to check this model out yet, how does it fare here? What's needed to improve the models in this area now that the voice part is more or less solved?
That, and also using English words in the middle of another language phrase confuses them a lot.