The music examples are utterly fascinating. It sounds insanely natural. The only thing I can hear that sounds unnatural, is the way that the reverberation in the room (the "echo") immediately gets lower when the raw piano sound itself gets lower. In a real room, if you produce a loud sound and immediately after a soft sound, the reverberation of the loud sound remains. But since this network only models "the sound ri…
I can hear some distortion in the piano notes - which may be an audio compression artefact, or it may be the output of the resynthesis process. If you train NNs at the phrase level and overfit, then you get something that is indeed more or less the same as cross-fading at random between short sections. Piano music is very idiomatic, so you'll capture some typical piano gestures that way. But I'd be surprised if the m…
WaveNet: A Generative Model for Raw Audio
81–90 of 153 posts
Re: WaveNet: A Generative Model for Raw Audio
#82This is incredible. I'd be worried if I were a professional audiobook reader :)
I wouldn't. The results they offer are excellent, but the missing points they need to achieve human level are related to producing the correct intonation, which requires accurate understanding of the material. That is still at least ten years in the future, I expect.
A big problem with generating prosody has always been that our theories of it don't really provide a great prediction of people's behaviours. It's also very expensive to get people to do the prosody annotations accurately, using whatever given theory.
Predicting the raw audio directly cuts out this problem. The "theory" of prosody can be left latent, rather than specified explicitly.
Re: WaveNet: A Generative Model for Raw Audio
#83Earlier quoted context omitted.
To my Australian English ears, the babbling sounded vaguely Scandinavian.
Indeed. I was surprised by that as well. Sounded like a Dutch speaker with a muffled voice behind a screen.
Re: WaveNet: A Generative Model for Raw Audio
#84Is it possible to use the "deep dream" methods with a network trained for audio such as this? I wonder what that would sound like, e.g., beginning with a speech signal and enhancing with a network trained for music or vice versa.
Re: WaveNet: A Generative Model for Raw Audio
#85The music examples are utterly fascinating. It sounds insanely natural. The only thing I can hear that sounds unnatural, is the way that the reverberation in the room (the "echo") immediately gets lower when the raw piano sound itself gets lower. In a real room, if you produce a loud sound and immediately after a soft sound, the reverberation of the loud sound remains. But since this network only models "the sound ri…
It shot me forward to a time where people just click a button to generate music they want to listen to. If you really like the generation, you save it and share it. It wouldn't have all of the other aspects that we derive from human-produced music like soul/emotion (because we know it's coming from a human, not because of how it sounds), but it would be a cool application idea anyway.
Re: WaveNet: A Generative Model for Raw Audio
#86Re: WaveNet: A Generative Model for Raw Audio
#87Earlier quoted context omitted.
I wouldn't. The results they offer are excellent, but the missing points they need to achieve human level are related to producing the correct intonation, which requires accurate understanding of the material. That is still at least ten years in the future, I expect.
Not really. They're training directly on the waveform, so the model can learn intonation. They just need to train on longer samples, and perhaps augment their linguistic representation with some extra discourse analysis. A big problem with generating prosody has always been that our theories of it don't really provide a great prediction of people's behaviours. It's also very expensive to get people to do the prosody…
Re: WaveNet: A Generative Model for Raw Audio
#88Is it possible to use the "deep dream" methods with a network trained for audio such as this? I wonder what that would sound like, e.g., beginning with a speech signal and enhancing with a network trained for music or vice versa.
We tried this but with less success than what wavenet did. https://wp.nyu.edu/ismir2016/wp-content/uploads/sites/2294/2...
Re: WaveNet: A Generative Model for Raw Audio
#89Earlier quoted context omitted.
Audio quality does leave something to be desired. https://vimeo.com/47987691
Lincoln died before Edison invented the phonograph. That's a hoax.
These early recordings are incredibly crude, and they did not have the technology at the time to play them back. They were just experiments in trying to view sound waves, not attempts to preserve information for future generations.
Re: WaveNet: A Generative Model for Raw Audio
#90Earlier quoted context omitted.
Not really. They're training directly on the waveform, so the model can learn intonation. They just need to train on longer samples, and perhaps augment their linguistic representation with some extra discourse analysis. A big problem with generating prosody has always been that our theories of it don't really provide a great prediction of people's behaviours. It's also very expensive to get people to do the prosody…
theres 0 chance of effective intonation and tone without understanding of the material