Live data from Hacker News

StreamVC: Real-Time Low-Latency Voice Conversion

research.google

11–20 of 44 posts

Re: StreamVC: Real-Time Low-Latency Voice Conversion

#13

What is the current best Foss(or otherwise) implementation for voice changer/anonymiser?

Last time I checked, it was https://github.com/w-okada/voice-changer Requires a decent amount of VRAM and runs poorly with pretty bad quality (IMO)

Once again we see evidence that AI-for-all is not bottlenecked by research but by the physical limitations of compute infrastructure.

Re: StreamVC: Real-Time Low-Latency Voice Conversion

#14
post #6
post #5

The samples were released a while back: https://google-research.github.io/seanet/stream_vc/

Not a very good demo page. It's difficult to judge real world quality with such unenthusiastic reading, unrealistic sentences, and unfamiliar voices. Typical of speech papers. It would be much better if celebrities were used as target voices, as we all know what they sound like and can therefore judge quality better. But I suppose that would be too controversial for Google. In general I think it is silly that voice c…

> But I suppose that would be too controversial for Google.

You don't have to suppose anything: it is actually settled law that its bad to just willy-nilly use people's voices if you feel like it, even if its just a sound-alike!

Re: StreamVC: Real-Time Low-Latency Voice Conversion

#15
post #13

Earlier quoted context omitted.

Last time I checked, it was https://github.com/w-okada/voice-changer Requires a decent amount of VRAM and runs poorly with pretty bad quality (IMO)

Once again we see evidence that AI-for-all is not bottlenecked by research but by the physical limitations of compute infrastructure.

I wouldn't say that when the application is a strung-up Python Frankenstein monster (not to be too demeaning to the author).

Re: StreamVC: Real-Time Low-Latency Voice Conversion

#16
post #13

Earlier quoted context omitted.

Last time I checked, it was https://github.com/w-okada/voice-changer Requires a decent amount of VRAM and runs poorly with pretty bad quality (IMO)

Once again we see evidence that AI-for-all is not bottlenecked by research but by the physical limitations of compute infrastructure.

More efficient architectures are possible. It's bottlenecked by research.

Re: StreamVC: Real-Time Low-Latency Voice Conversion

#17
post #5

The samples were released a while back: https://google-research.github.io/seanet/stream_vc/

For those confused as I was - it's not trying to match the accent of the target speech in those samples, just the timbre. To quote the paper:

> Voice conversion refers to altering the style of a speech signal while preserving its linguistic content. While style encompasses many aspects of speech, such as emotion, prosody, accent, and whispering, in this work we focus on the conversion of speaker timbre only while keeping the linguistic and para-linguistic information unchanged.

Re: StreamVC: Real-Time Low-Latency Voice Conversion

#20
post #6

Earlier quoted context omitted.

Not a very good demo page. It's difficult to judge real world quality with such unenthusiastic reading, unrealistic sentences, and unfamiliar voices. Typical of speech papers. It would be much better if celebrities were used as target voices, as we all know what they sound like and can therefore judge quality better. But I suppose that would be too controversial for Google. In general I think it is silly that voice c…

> But I suppose that would be too controversial for Google. You don't have to suppose anything: it is actually settled law that its bad to just willy-nilly use people's voices if you feel like it, even if its just a sound-alike!

So, what do we do with actual people who have a very similar voice to some "more famous" person? It's quite silly when voices are far away from being unique to a person.
Post reply on HN