Earlier quoted context omitted.
With timecode metadata and heatmaps to highlight rapidfire back 'n forth and stuff? Semantic scoring to rate quip chains Vs. slower, longer considered interactions? Ways of highlighting high frequency commenters, perhaps rating frequent flyers by degree of interest in what they have to say and weighting threads by "participant rating". There are many ways to go in this space.
The best idea I have now is to scan HN and make a list of commenters who said something like “I am working on a PhD in this topic”.
There are a couple of angles, straight up representation of forum | subreddit threads after the conversation has moved on, "live" tracking comments as conversations progress, and moderation of live threads (including swatting | detecting trolls, spambots, griefers, etc) (oh, and retro sweeps looking for tail end spammers and their puppet networks eg: https://news.ycombinator.com/item?id=40800160 (with "ShowDead" on)).
In post analysis there's the whole ball of wax around decrufting chat threads to pull the meat best served to (various) AI, either for training, company | group intelligence, etc. Oh, and fingerprinting across multiple sources looking for common {group | individual} traits
The intrinsic issue with excluding all but those with a current PhD on chat would be the tossing of, say, those with dated PhD in theoretical physics or somesuch.
@dang obviously mods here and likely has a slew of lisp-y scripts | tools for giving various views, @jedberg was about for much of the early evolution of reddit and took part in a lot of discussions about presentation and implemented a few.
There'll also be useful input from any of those who've moderated | admin various largish forums | channels, etc over the years.