I'm sure I'm missing something, but: aren't there publicly available corpuses of all reddit posts up to a certain date? Why wouldn't researchers train with these? Are they just not recent enough? Even if they aren't very recent, how big of a deal is that when it comes to training models that presumably use lots of other data sources as well?
> Most of Reddit’s current data has been scraped anyway, so the game is to protect Reddit’s data going forward.
But yes, Pushshift archives of all posts and comments until the recent ban [1] are freely available for download [2]
[1] https://old.reddit.com/r/pushshift/comments/135tdl2/a_respon... The ban was followed by allowing the parent non-profit of Pushshift (Network Contagion Research Institute) to use the API provided access is restricted to a use-case Reddit has care for: mod tools. Reddit hasnt replaced those with its own just yet. The rest of us are shut out.