For those too impatient to read the details, check out the "Hear for yourself" examples toward the bottom of the page. They're reproducing decent sounding speech at 1.6 kbps. 1.6 kbps is nuts! I like to re-encode audio books or podcasts in Opus at 32 kbps and I consider that stingy. The fact that speech is even comprehensible at 1.6 kbps is impressive. As the article explains, their technique is analogous to speech-t…
Actually, this won't work at all for music because it makes fundamental assumptions that the signal is speech. For normal conversations, it should work, though for now the models are not yet as robust as I'd like (in case of noise and reverberation). That's next on the list of things to improve.
A Real-Time Wideband Neural Vocoder at 1.6 Kb/S Using LPCNet
11–20 of 76 posts
Re: A Real-Time Wideband Neural Vocoder at 1.6 Kb/S Using LPCNet
#123 Gflops, we are deep beyond diminishing returns here. Opus seems good enough.
Re: A Real-Time Wideband Neural Vocoder at 1.6 Kb/S Using LPCNet
#13For those too impatient to read the details, check out the "Hear for yourself" examples toward the bottom of the page. They're reproducing decent sounding speech at 1.6 kbps. 1.6 kbps is nuts! I like to re-encode audio books or podcasts in Opus at 32 kbps and I consider that stingy. The fact that speech is even comprehensible at 1.6 kbps is impressive. As the article explains, their technique is analogous to speech-t…
https://en.wikipedia.org/wiki/Adaptive_Multi-Rate_audio_code...
Re: A Real-Time Wideband Neural Vocoder at 1.6 Kb/S Using LPCNet
#143 Gflops, we are deep beyond diminishing returns here. Opus seems good enough.
Opus isn't good enough to be a replacement for AMBE for use over radio. Opus doesn't make it easier to make very high quality speech synthesis, etc.
Opus loss robustness could be much better using tools from this toolbox-- and we're a long way from not wanting better performance in the face of packet loss.
Re: A Real-Time Wideband Neural Vocoder at 1.6 Kb/S Using LPCNet
#15For those too impatient to read the details, check out the "Hear for yourself" examples toward the bottom of the page. They're reproducing decent sounding speech at 1.6 kbps. 1.6 kbps is nuts! I like to re-encode audio books or podcasts in Opus at 32 kbps and I consider that stingy. The fact that speech is even comprehensible at 1.6 kbps is impressive. As the article explains, their technique is analogous to speech-t…
Actually, this won't work at all for music because it makes fundamental assumptions that the signal is speech. For normal conversations, it should work, though for now the models are not yet as robust as I'd like (in case of noise and reverberation). That's next on the list of things to improve.
Re: A Real-Time Wideband Neural Vocoder at 1.6 Kb/S Using LPCNet
#16For those too impatient to read the details, check out the "Hear for yourself" examples toward the bottom of the page. They're reproducing decent sounding speech at 1.6 kbps. 1.6 kbps is nuts! I like to re-encode audio books or podcasts in Opus at 32 kbps and I consider that stingy. The fact that speech is even comprehensible at 1.6 kbps is impressive. As the article explains, their technique is analogous to speech-t…
Re: A Real-Time Wideband Neural Vocoder at 1.6 Kb/S Using LPCNet
#17Earlier quoted context omitted.
Actually, this won't work at all for music because it makes fundamental assumptions that the signal is speech. For normal conversations, it should work, though for now the models are not yet as robust as I'd like (in case of noise and reverberation). That's next on the list of things to improve.
Here we go! This is the first minute or so of Penny Lane by The Beatles converted down to a 10KB .bin and then back to a .wav: http://no.gd/pennylane.wav .. unsurprisingly the vocals remain recognizable, but the music barely at all.
Re: A Real-Time Wideband Neural Vocoder at 1.6 Kb/S Using LPCNet
#18Re: A Real-Time Wideband Neural Vocoder at 1.6 Kb/S Using LPCNet
#19Earlier quoted context omitted.
Here we go! This is the first minute or so of Penny Lane by The Beatles converted down to a 10KB .bin and then back to a .wav: http://no.gd/pennylane.wav .. unsurprisingly the vocals remain recognizable, but the music barely at all.
As imagined by Marilyn Manson...
I've also run a BBC news report through the program with better results although it demonstrates that any background noise at all can throw things off significantly: https://twitter.com/peterc/status/1111736029558517760 .. so at this low bitrate, it really is only good for plain speech without any other noise.