For anyone interested in the Raw Audio Models section of that articles, there are some fun endpoints[1][2] from the Echonest API that provide those models. They've moved over to the Spotify API since I last had a play, but it's great that they still provide them. You can get the audio breakdown of a track, as well as a summary of the track features including fun stuff like "danceability" and musical positiveness ("va…
The neural network is trained to mimic collaborative filtering vectors from raw audio. It's a separate model from time signature, key, mode, tempo, loudness, etc.