Live data from Hacker News

Inflect-Micro-v2: complete voice in 9.36M parameters

huggingface.co

1–10 of 35 posts

Re: Inflect-Micro-v2: complete voice in 9.36M parameters

#4
Couple highlights:

> Complete local text-to-waveform speech synthesis under 10M parameters.

In case, like me, you hoped "complete" voice might mean both stt and tts. Not to speak poorly of it, just clarifying.

> English only, with one fixed male voice. This is not zero-shot voice cloning.

(And then a bunch of statements on limitations that I read as 'quality can be spotty but if you play with it it should be fine') But like. In <10M params I'm not judging:)

Re: Inflect-Micro-v2: complete voice in 9.36M parameters

#9
post #8

Amazing quality for small size, but definitely not that enjoyable to listen to. IMHO, its at about the same quality level of historic TTS tools.

I'm not sure which historic tools you mean, but to me this sounds much better than anything older than ten years ago.
Post reply on HN