Live data from Hacker News

Million Song Dataset

millionsongdataset.com

41–44 of 44 posts

Re: Million Song Dataset

#41
post #36
post #5

I think this Dataset is probably unfortunately most famous for an incredibly flawed but very headlineable attempt, which you've almost certainly seen somewhere, by a group of researchers from AI and related fields (none of which had musical qualifications) to "objectively" determine if music has gotten worse over the decades. As usual, they arrived at the conclusion that it did by computing a vague number ("timbral d…

I wonder if digital recording and processing has made music cleaner leading to less harmonic garbage. Also... What if the poetry is better, and someone is singing acapella? How do you capture that?

Yeah, this is really the core issue, even within the music.

As a thought experiment, let's say someone invented a brilliant revolutionary new complex rhythm. Everyone went wild using it for a year. Then, once everyone knows it, artists start to mix it up by leaving out parts of it, leaving it implied, relying on listeners familiarity with the rhythm to make things work.

If you now tried to naively measure the amount of rhythmic complexity per song by counting percussion hits or similar, you'd see complexity take a nosedive. You'd also see people who missed out on the year complain about how bland the new rhythms are. But the songs actually got more complex. It's just that the complexity is only apparent to people familiar with the hypothetical revolutionary rhythm.

At the same time, people familiar with the new music will look back at the old music and be incredibly bored. They're used to finding enjoyment in the complexity of the implied rhythm, but there's just nothing there, it's all painfully spelled out and predictable.

Re: Million Song Dataset

#42

Earlier quoted context omitted.

No, I think the poster is bothered by how absurd and unscientific such a pursuit is. It proposes that perhaps the people who made this dataset don't actually know anything about music, and are just nerds. This behavior is a common form of malpractice in composing data sets such as these.

"Worse" might be subjective but if they analyzed the data and found decreasing complexity of harmony and song structure, then that's not what I would call "unscientific". For the past 30 years music has definitely become more "bland" (for lack of a better word). The data backs this up. Whether you like that or not is subjective.

I'm actually kind of going to agree with you.

First of all, I don't like the term "bland", because it implies a judgement. I think a better word is "sparse".

But now, the question, and the thing the researchers got fatally wrong is: does sparseness really mean less complexity? When you phrase it like that, it sounds obviously silly, to me at least. You wouldn't judge say Picasso for not using enough colors, and complain about it being so bland. To continue on that path, Picasso only works because you know what things actually look like. Instead of having to painfully spell out every last facial hair, he relies on you knowing what a human looks like, and uses the space gained to express his intent more clearly. In comparison, older art might seem boring and predictable. I already know what a face looks like, why waste space showing me?

Modern music works on a similar principle. The sparseness of modern music is enabled by listeners familiarity with other music, which enables composers to simply sufficiently imply their intent, letting our brains fill in the rest, making space to innovate in other areas of the music.

Re: Million Song Dataset

#43
post #42

Earlier quoted context omitted.

"Worse" might be subjective but if they analyzed the data and found decreasing complexity of harmony and song structure, then that's not what I would call "unscientific". For the past 30 years music has definitely become more "bland" (for lack of a better word). The data backs this up. Whether you like that or not is subjective.

I'm actually kind of going to agree with you. First of all, I don't like the term "bland", because it implies a judgement. I think a better word is "sparse". But now, the question, and the thing the researchers got fatally wrong is: does sparseness really mean less complexity? When you phrase it like that, it sounds obviously silly, to me at least. You wouldn't judge say Picasso for not using enough colors, and compl…

Very interesting. I probably should have clarified that I was talking about popular music and not niche genres. I actually think the effect is often less about sparseness bringing the composer's vision to light and more to do with the trend toward group songwriting.

For instance Beyonce had 72 songwriters [1] listed on her Lemonade album. Compare that with Madonna [2] (and other contemporaries) who often write their own music or co-write a song with one other person.

1. https://www.thedailybeast.com/does-beyonce-write-her-own-mus...

2. https://www.quora.com/How-much-of-her-music-does-Madonna-act...

Re: Million Song Dataset

#44
post #28

Earlier quoted context omitted.

For virtually any kind of art newer artworks will be on average worse than _surviving_ older artworks: the older artworks have undergone a selection process that weeded out the less popular ones. This is IMO a more fundamental problem with any comparisons across long time intervals.

True for a naive approach, but I'd think looking at the chart toppers for, say, each week would give decently comparable data.

Yes, assuming you can find all the old chart toppers (and that this term has a consistent meaning over the whole time interval you're looking at). If, understandably, you have some gaps in old chart toppers, you need to assume something about them (probably that they were worse than all surviving ones, which helps if you're doing percentile-based comparisons).
Post reply on HN