Live data from Hacker News

Finding the genre of a song with Deep Learning

medium.com

21–30 of 36 posts

Re: Finding the genre of a song with Deep Learning

#21
post #5
post #2

Wow, I find it incredible that this works. As I understand it, the approach is to do a Fourier transform on a couple seconds of the song to create a 128x128 pixel spectrogram. Each horizontal pixel represents a 20 ms slice in time, and each vertical pixel represents 1/128 of the frequency domain. Then treating these spectrograms as images, train a neural net to classify them using pre-labelled samples. Then take samp…

One reason might be that the mentioned genres are highly formulaic to begin with. The standard rap song contains about 2 bars of unique music stretched out over 3 minutes with slight variations. Same with dubstep and techno. All highly repetitive. Classical music got no drums, so you can detect that. Metal got guitar distortion all over the spectrum. So with these examples the spectral images should have enough disti…

It's true that having very different genres helps the model a lot. It would be much more difficult to distinguish between closer genres, especially when people don't really know which is which and argue all the time about it.

Re: Finding the genre of a song with Deep Learning

#22
post #3

Nice approach, and well explained! By the way, Niland is a startup that also does music labeling with the help of deep learning. Demo available here: http://demo.niland.io/ For example, it can output Drum Machine: 87%, House: 88%, Female Voice: 55%, Groovy: 93%

Thanks for the kind words, I'll take a look !

Re: Finding the genre of a song with Deep Learning

#23

That's pretty cool, I'd like to use something like this to tell me what genre my own songs are. It's annoying to write a song and then upload it to some service or another and have no idea what genre to pick. :-) My stuff is somewhere in the jazz-influenced singer-songwriter american piano pop realm which is a combination that works for me but it generally feels like I'm selling the song short if I have to pick only…

Yeah that's a problem I know - I used to make some Electro/Dubstep/Trap music - and I feel people will always disagree with the genre you pick anyway.

Re: Finding the genre of a song with Deep Learning

#24
post #12

Hmm, convolution is perfectly good operation to run on wave forms as well. In fact the wikipedia article ( https://en.wikipedia.org/wiki/Convolution ) shows the operation on functions which would correspond to time-domain wave forms. What is the point of converting everything to pictures and then using 2D convolutions when that step could have been skipped entirely? Converting to pictures is unnecessary. It makes the…

The idea is that the vertical axis of the spectrogram is basically already an hierarchical set of features (in scale/frequency). Then convolutions on that is a lot like how DenseNets combine hierarchical features. I agree it seems a little jank, but the features are pretty good - and a lot of network architectures / training techniques are most practiced in an image processing context.

Thanks for your inputs, it's true that we can use convolutions on raw waveform, however the main reason I've used a spectrogram was to work on precomputed relevant features as highd pointed out, instead of running the convolution on lots of data.

Re: Finding the genre of a song with Deep Learning

#26
post #4

To the author: Have you tried to use a logarithmic frequency scale in the spectrogram? [1] That representation is closer to the way humans perceive sound, and gives you finer resolution in the lower frequencies. [2] If you want to make your representation even closer to the human's perception, take a look at Google's CARFAC research. [3] Basically, they model the ear. I've prepared a Python utility for converting sou…

I don't think this problem is bound by absolute frequency resolution, the tightest distance between two notes on a typical piano is ~2hz and if you assume a doubling between octaves you're at <90 notes. The temporal changes and relative chord progressions probably give more info.

Thanks for your insights! I agree that log/mel spectrograms could be even more detailed and effective, and could be used with the SoX patch discussed here https://sourceforge.net/p/sox/feature-requests/176/.

Re: Finding the genre of a song with Deep Learning

#27
post #4

To the author: Have you tried to use a logarithmic frequency scale in the spectrogram? [1] That representation is closer to the way humans perceive sound, and gives you finer resolution in the lower frequencies. [2] If you want to make your representation even closer to the human's perception, take a look at Google's CARFAC research. [3] Basically, they model the ear. I've prepared a Python utility for converting sou…

I didn't intend to go that far in the human genre recognition parallel, but thanks for the references ! Good job on the script too

Re: Finding the genre of a song with Deep Learning

#28
1. I wonder how the continuous wavelet transform would compare to the windowed Fourier transform used here. See [1] an python implementation, for example.

2. The size of frequency analysis blocks seems arbitrary. I wonder if there is a "natural" block size based on a song's tempo, say 1 bar. This would of course require a priori tempo knowledge or a run-time estimate.

[1]: https://docs.scipy.org/doc/scipy-0.15.1/reference/generated/...

Re: Finding the genre of a song with Deep Learning

#29

1. I wonder how the continuous wavelet transform would compare to the windowed Fourier transform used here. See [1] an python implementation, for example. 2. The size of frequency analysis blocks seems arbitrary. I wonder if there is a "natural" block size based on a song's tempo, say 1 bar. This would of course require a priori tempo knowledge or a run-time estimate. [1]: https://docs.scipy.org/doc/scipy-0.15.1/refe…

The slice size is indeed quite arbitrary, and knowing the BPM would help, but isn't reliable either (various tempos, rubato for classical etc.)

Re: Finding the genre of a song with Deep Learning

#30
post #2

Wow, I find it incredible that this works. As I understand it, the approach is to do a Fourier transform on a couple seconds of the song to create a 128x128 pixel spectrogram. Each horizontal pixel represents a 20 ms slice in time, and each vertical pixel represents 1/128 of the frequency domain. Then treating these spectrograms as images, train a neural net to classify them using pre-labelled samples. Then take samp…

I guess the spectogram behaves like an image in that translation of any feature by an arbitrary distance (dx,dy) preserves its predicting quality.

But please correct me if I'm wrong.

Post reply on HN