Live data from Hacker News

Million Song Dataset

millionsongdataset.com

31–40 of 44 posts

Re: Million Song Dataset

#31
post #28
post #5

I think this Dataset is probably unfortunately most famous for an incredibly flawed but very headlineable attempt, which you've almost certainly seen somewhere, by a group of researchers from AI and related fields (none of which had musical qualifications) to "objectively" determine if music has gotten worse over the decades. As usual, they arrived at the conclusion that it did by computing a vague number ("timbral d…

For virtually any kind of art newer artworks will be on average worse than _surviving_ older artworks: the older artworks have undergone a selection process that weeded out the less popular ones. This is IMO a more fundamental problem with any comparisons across long time intervals.

True for a naive approach, but I'd think looking at the chart toppers for, say, each week would give decently comparable data.

Re: Million Song Dataset

#32

Earlier quoted context omitted.

"Worse" might be subjective but if they analyzed the data and found decreasing complexity of harmony and song structure, then that's not what I would call "unscientific". For the past 30 years music has definitely become more "bland" (for lack of a better word). The data backs this up. Whether you like that or not is subjective.

It's tempting to think one can come up with objective measures of musical complexity, but the problem is the number of dimensions musical complexity exists in: Many people who listen to classical (in the more general sense of the word that encompasses baroque and romantic eras as well) dislike folk music as being too simple. They are listening to harmonic complexity - the changing of chords over time. Irish music usu…

Very well put!

Re: Million Song Dataset

#33

Does this data set have a collection of chords in text? I'd love something like it.

The dataset is a bunch of SQLite files, so shouldn't be too tricky to interrogate. You can get a subset (static.echonest.com/millionsongsubset_full.tar.gz), which is 1.8Gb compressed. The full dataset is 280Gb, and AFAICT, this does not contain the full audio. There's a script on GitHub from like 8 years ago that apparently can get you the audio (but I would be super-impressed if that actually still works). You can s…

It looks like they have data about many of the individual notes in the song. I wonder if it could be possible to turn that data back into some sort of horrible midi version.

Re: Million Song Dataset

#34
post #28

Earlier quoted context omitted.

For virtually any kind of art newer artworks will be on average worse than _surviving_ older artworks: the older artworks have undergone a selection process that weeded out the less popular ones. This is IMO a more fundamental problem with any comparisons across long time intervals.

True for a naive approach, but I'd think looking at the chart toppers for, say, each week would give decently comparable data.

I actually did something similar! I wrote a small blog post about it here https://www.popnalysis.com/blog/lyrics-over-time/

But I came to the same conclusion as the op comment. Music can't really be judged by any one metric! But that doesn't mean that you don't gain insight into how music has evolved!

also, I did release my entire lyrics dataset that I scraped (about 500k) for free.

Re: Million Song Dataset

#35
post #28
post #5

I think this Dataset is probably unfortunately most famous for an incredibly flawed but very headlineable attempt, which you've almost certainly seen somewhere, by a group of researchers from AI and related fields (none of which had musical qualifications) to "objectively" determine if music has gotten worse over the decades. As usual, they arrived at the conclusion that it did by computing a vague number ("timbral d…

For virtually any kind of art newer artworks will be on average worse than _surviving_ older artworks: the older artworks have undergone a selection process that weeded out the less popular ones. This is IMO a more fundamental problem with any comparisons across long time intervals.

Selection bias is also apparent in ranking sites for shows where seasons/sequels are ranked individually.

For example in anime, Gintama appears 8 times within the top 50: https://myanimelist.net/topanime.php

It's not because it's a popular show. It's just that people who didn't like the first few episodes have already stopped watching! And it's polarizing enough that the only people who stick around for so many seasons are the ones really love it. So it will get rated a 10 even for mediocre content.

Re: Million Song Dataset

#36
post #5

I think this Dataset is probably unfortunately most famous for an incredibly flawed but very headlineable attempt, which you've almost certainly seen somewhere, by a group of researchers from AI and related fields (none of which had musical qualifications) to "objectively" determine if music has gotten worse over the decades. As usual, they arrived at the conclusion that it did by computing a vague number ("timbral d…

I wonder if digital recording and processing has made music cleaner leading to less harmonic garbage.

Also... What if the poetry is better, and someone is singing acapella? How do you capture that?

Re: Million Song Dataset

#37
For anyone else who was curious about the song selection process:

How did you choose the million tracks?

Choosing a million songs is surprisingly challenging. We followed these steps:

1. Getting the most 'familiar' artists according to The Echo Nest, then downloading as many songs as possible from each of them 2. Getting the 200 top terms from The Echo Nest, then using each term as a descriptor to find 100 artists, then downloading as many of their songs as possible 3. Getting the songs and artists from the CAL500 dataset 4. Getting 'extreme' songs from The Echo Nest search params, e.g. songs with highest energy, lowest energy, tempo, song hotttnesss, ... 5. A random walk along the similar artists links starting from the 100 most familiar artists

The number of songs was approximately 8950 after step 1), step 3) added around 15000 songs, and we add approx. 500000 songs before starting step 5. For more technical details, see "dataset creation" in the "code" tab.

[1]

What I really wanted to know was if it was a worldwide-music dataset or a more narrowly focused one. My guess based on the above is that it's mostly English-language music, mostly American - can someone who's worked with the data confirm/deny that?

[1] http://millionsongdataset.com/faq/#how-did-you-choose-millio...

Re: Million Song Dataset

#39

Earlier quoted context omitted.

"Worse" might be subjective but if they analyzed the data and found decreasing complexity of harmony and song structure, then that's not what I would call "unscientific". For the past 30 years music has definitely become more "bland" (for lack of a better word). The data backs this up. Whether you like that or not is subjective.

This is such nonsense it's just incredible that people with such seemingly strong conversational strengths in music could say something so perversely false. If you are looking at the kind of music that makes it onto Top 40 radio stations, then this is undoubtedly true. But that's not what is meant when you say > For the past 30 years music has definitely become more "bland" When you suggest that, you appear naive and…

There has always been underground music and there always will be (thankfully). I'm talking about mainstream music, which most definitely has become simplified in the past 30 years.

Re: Million Song Dataset

#40
post #23

Earlier quoted context omitted.

"Worse" might be subjective but if they analyzed the data and found decreasing complexity of harmony and song structure, then that's not what I would call "unscientific". For the past 30 years music has definitely become more "bland" (for lack of a better word). The data backs this up. Whether you like that or not is subjective.

If all you listen to is the radio, sure. However in the broader sense of music being created, this is patently false. There are more genres and subgenres of music than ever before, experimentation and synthesis techniques weaving together ever more intricate and sophisticated timbres. Movies are a better example of a medium getting bland, but there is a material reason for that; production costs vs return on investme…

I am talking about popular music since the vast majority of people do only listen to the radio.

Of course there's a lot of underground and independent music being made in many genres. It just doesn't have much of an audience.

Post reply on HN