StreamVC: Real-Time Low-Latency Voice Conversion
11–20 of 44 posts
Re: StreamVC: Real-Time Low-Latency Voice Conversion
#12What is the current best Foss(or otherwise) implementation for voice changer/anonymiser?
Requires a decent amount of VRAM and runs poorly with pretty bad quality (IMO)
Re: StreamVC: Real-Time Low-Latency Voice Conversion
#13What is the current best Foss(or otherwise) implementation for voice changer/anonymiser?
Last time I checked, it was https://github.com/w-okada/voice-changer Requires a decent amount of VRAM and runs poorly with pretty bad quality (IMO)
Re: StreamVC: Real-Time Low-Latency Voice Conversion
#14The samples were released a while back: https://google-research.github.io/seanet/stream_vc/
Not a very good demo page. It's difficult to judge real world quality with such unenthusiastic reading, unrealistic sentences, and unfamiliar voices. Typical of speech papers. It would be much better if celebrities were used as target voices, as we all know what they sound like and can therefore judge quality better. But I suppose that would be too controversial for Google. In general I think it is silly that voice c…
You don't have to suppose anything: it is actually settled law that its bad to just willy-nilly use people's voices if you feel like it, even if its just a sound-alike!
Re: StreamVC: Real-Time Low-Latency Voice Conversion
#15Earlier quoted context omitted.
Last time I checked, it was https://github.com/w-okada/voice-changer Requires a decent amount of VRAM and runs poorly with pretty bad quality (IMO)
Once again we see evidence that AI-for-all is not bottlenecked by research but by the physical limitations of compute infrastructure.
Re: StreamVC: Real-Time Low-Latency Voice Conversion
#16Earlier quoted context omitted.
Last time I checked, it was https://github.com/w-okada/voice-changer Requires a decent amount of VRAM and runs poorly with pretty bad quality (IMO)
Once again we see evidence that AI-for-all is not bottlenecked by research but by the physical limitations of compute infrastructure.
Re: StreamVC: Real-Time Low-Latency Voice Conversion
#17The samples were released a while back: https://google-research.github.io/seanet/stream_vc/
> Voice conversion refers to altering the style of a speech signal while preserving its linguistic content. While style encompasses many aspects of speech, such as emotion, prosody, accent, and whispering, in this work we focus on the conversion of speaker timbre only while keeping the linguistic and para-linguistic information unchanged.
Re: StreamVC: Real-Time Low-Latency Voice Conversion
#18Re: StreamVC: Real-Time Low-Latency Voice Conversion
#19https://github.com/hrnoh24/stream-vc https://github.com/yuval-reshef/StreamVC Unofficial implementations of StreamVC
Re: StreamVC: Real-Time Low-Latency Voice Conversion
#20Earlier quoted context omitted.
Not a very good demo page. It's difficult to judge real world quality with such unenthusiastic reading, unrealistic sentences, and unfamiliar voices. Typical of speech papers. It would be much better if celebrities were used as target voices, as we all know what they sound like and can therefore judge quality better. But I suppose that would be too controversial for Google. In general I think it is silly that voice c…
> But I suppose that would be too controversial for Google. You don't have to suppose anything: it is actually settled law that its bad to just willy-nilly use people's voices if you feel like it, even if its just a sound-alike!