Now that's an interesting regression! I don't remember seeing it there before.
(The worst I've noticed before has been Lawrence of Arabia under History of the Ancient World. Very much 20th century really.)
Several other classifications are arguable -- which I think shows one of the limitations of this technique: it's not possible to iterate + improve.
So instead I've been wondering about using the embeddings of each episode synopsis, and comparing to the embeddings of Dewey subdivisions. I should be able to tune the results better that way.
There's also a technique from Google called CAVs (Concept Activation Vectors) that I'm intrigued about trying -- would love to hear if anybody has experience using this
https://arxiv.org/abs/1711.11279