I commented on this the last time it made the rounds, and the study's methodology is flawed in a pretty central way.
From the paper:
> Hate speech and harassment are contentious
topics, lacking clear definitions. As discussed above, the European Court of Human Rights notes that
“no universally accepted definition of the term ‘hate speech’ exists” [48]. They adopt a definition
including “comments which are necessarily directed against a person or particular group of people”,
focusing on race, religion, “aggressive nationalism and ethnocentrism”, and homophobic speech [48].
> This definition provides a useful starting point, but it is difficult to operationalize at scale.
We therefore take a usage-based approach: given that Reddit has banned the r/fatpeoplehate
and r/CoonTown forums, we focus on textual content that is distinctively characteristic of these
forums. Using an automated keyword identification technique, we build lexicons of keywords for
r/fatpeoplehate and r/CoonTown, which makes it possible to track whether the words in these
lexicons become more common in other forums after the ban. Next, we manually inspect the
automatically generated lexicons, and identify a subset of terms that are especially oriented towards
hate speech. These manually refined lexicons are sparser, but offer higher precision.
The Tldr here is that they defined hate speech as "the subset of terms heavily used in these subreddits, manually filtered by the authors' subjective (and opaque) assessment of hatefulness". My beef actually isn't with the subjectivity of the latter part: I think the reliance on the ill-defined concept of hate speech does a lot more harm than good, but I acknowledge that plenty of people think it's valuable, and that's a non-central hill to die on here.
The real issue with the study is that starting your definition of hate speech with "the lexicon of banned subs" is a fatal confounder:
Many subreddits, particularly free-flowing ideological echo chambers, end up with distinctive lexicons: specific phrases, terms, and in-jokes that people bring up to solidify feelings of community. This isn't a novel insight, since it's how most communities (and indeed relationships) work. the phenomenon is exacerbated for ideological echo chambers, and not just the ones in the right which are more likely to align with people's definition of hate speech: you can find the exact same phenomenon on exho chambers like LateStageCapitalism or ChoosingBeggars, and a milder version on most subs that are less general-interest than eg r/movies.
The study's conclusions essentially boils down to: "if you ban a sub, the distinctive lexicon of that sub becomes less common".....duh? It provides very little signal about the way the amount of hate speech evolves, as someone from fatpeoplehate could easily be spewing the same vile content elsewhere, and just use the term "hambeast" less in favor of a more broadly-used slur like "pig". This isn't a possibility that the paper even attempts to address, and it's conclusions in light of this fatal flaw are downright dishonest.
It's depressing how often people cite this study uncritically. As always, _read the papers behind articles before you cite or share believe them_, particularly when they're reliant on undefined, impossible to measure definitions like "hate speech". If authors at "papers of record" like the NYT are regularly too dumb or dishonest to accurately describes papers' conclusions (or at least describe their limitations), a rag like TechCrunch is DEFINITELY not something you should take at face value.