Web version: https://clowerweb.github.io/kitten-tts-web-demo/ It sounds ok, but impressive for the size.
Show HN: Kitten TTS – 25MB CPU-Only, Open-Source TTS Model
101–110 of 383 posts
Re: Show HN: Kitten TTS – 25MB CPU-Only, Open-Source TTS Model
#102The headline feature isn’t the 25 MB footprint alone. It’s that KittenTTS is Apache-2.0. That combo means you can embed a fully offline voice in Pi Zero-class hardware or even battery-powered toys without worrying about GPUs, cloud calls, or restrictive licenses. In one stroke it turns voice everywhere from a hardware/licensing problem into a packaging problem. Quality tweaks can come later; unlocking that deployment…
Re: Show HN: Kitten TTS – 25MB CPU-Only, Open-Source TTS Model
#103 System Requirements
Works literally everywhere
Haha, on one of my machines my python version is too old, and the package/dependencies don't want to install.On another machie the python version is too new, and the package/dependencies don't want to install.
Re: Show HN: Kitten TTS – 25MB CPU-Only, Open-Source TTS Model
#104Kudos guys!
Re: Show HN: Kitten TTS – 25MB CPU-Only, Open-Source TTS Model
#105Can you run it in reverse for speech recognition?
Re: Show HN: Kitten TTS – 25MB CPU-Only, Open-Source TTS Model
#106it would be great if there is typescript support in the future
Re: Show HN: Kitten TTS – 25MB CPU-Only, Open-Source TTS Model
#107Cool. While I think this is indeed impressive and has a specific use case (e.g. in the embedded sector), I'm not totally convinced that the quality is good enough to replace bigger models. With fish-speech[1] and f5-tts[2] there are at least 2 open source models pushing the quality limits of offline text-to-speech. I tested F5-TTS with an old NVidia 1660 (6GB VRAM) and it worked ok-ish, so running it on a little more…
Re: Show HN: Kitten TTS – 25MB CPU-Only, Open-Source TTS Model
#108I don't mind so much the size in MB, the fact that it's pure CPU and the quality, what I do mind however is the latency. I hope it's fast. Aside: Are there any models for understanding voice to text, fully offline, without training? I will be very impressed when we will be able to have a conversation with an AI at a natural rate and not "probe, space, response"
Any idea what factors play into latency in TTS models?
Re: Show HN: Kitten TTS – 25MB CPU-Only, Open-Source TTS Model
#109Re: Show HN: Kitten TTS – 25MB CPU-Only, Open-Source TTS Model
#110Earlier quoted context omitted.
A robotic, explicitly non-natural voice would be perfectly acceptable, and even desirable, in many situations[...]it'd be enough is the speech if easy to recognize. We've had formant synths for several decades, and they're perfectly understandable and require a tiny amount of computing power, but people tend not to want to listen to them: https://en.wikipedia.org/wiki/Software_Automatic_Mouth https://simulationcorner…
Well, this one is a bit too jarring to the ears.