Live data from Hacker News

Inflect-Micro-v2: complete voice in 9.36M parameters

huggingface.co

21–30 of 35 posts

Re: Inflect-Micro-v2: complete voice in 9.36M parameters

#21

Couple highlights: > Complete local text-to-waveform speech synthesis under 10M parameters. In case, like me, you hoped "complete" voice might mean both stt and tts. Not to speak poorly of it, just clarifying. > English only, with one fixed male voice. This is not zero-shot voice cloning. (And then a bunch of statements on limitations that I read as 'quality can be spotty but if you play with it it should be fine') B…

When would “text-to-waveform speech synthesis” ever imply speech to text?

Re: Inflect-Micro-v2: complete voice in 9.36M parameters

#23

I'd love to hear it but it seems your quota is exhausted.

wrote a web-demo, runs in your browser https://inflect-tts.geronimo-labs.com code: https://github.com/geronimi73/inflect-tts

Nice! Strangely, "Nano" sounds a lot better than "Micro" to me.

On my iPhone 14 Pro the page crashes after 2-3 plays. I wonder if it uses too much memory?

Re: Inflect-Micro-v2: complete voice in 9.36M parameters

#27

Earlier quoted context omitted.

wrote a web-demo, runs in your browser https://inflect-tts.geronimo-labs.com code: https://github.com/geronimi73/inflect-tts

Nice! Strangely, "Nano" sounds a lot better than "Micro" to me. On my iPhone 14 Pro the page crashes after 2-3 plays. I wonder if it uses too much memory?

right. happens on my 13 too. memory leak confirmed, not sure yet what's causing it

Re: Inflect-Micro-v2: complete voice in 9.36M parameters

#28

I keep seeing tts stories here. Is it just an interesting subset of the llm world, or is there a huge use case I’m somehow missing?

I use STT/TTS to interface with a local LLM for Home Assistant in my house.

I'm do so as well, i tried qwen3 omni 3 but it was ridiculously stupid, and i ended up with stt thinker and tts. kokoro for now

Re: Inflect-Micro-v2: complete voice in 9.36M parameters

#29

Couple highlights: > Complete local text-to-waveform speech synthesis under 10M parameters. In case, like me, you hoped "complete" voice might mean both stt and tts. Not to speak poorly of it, just clarifying. > English only, with one fixed male voice. This is not zero-shot voice cloning. (And then a bunch of statements on limitations that I read as 'quality can be spotty but if you play with it it should be fine') B…

When would “text-to-waveform speech synthesis” ever imply speech to text?

The HN title is "Inflect-Micro-v2: complete voice in 9.36M parameters".
Post reply on HN