A dataset of every Reddit comment
11–20 of 96 posts
Re: A dataset of every Reddit comment
#12There was also someone that created a dump of all HN data a while ago: https://github.com/sytelus/HackerNewsData
Re: A dataset of every Reddit comment
#13Anyone know how big it is yet?
Re: A dataset of every Reddit comment
#14Anyone know how big it is yet?
One bz2 file of comments per month.
Re: A dataset of every Reddit comment
#15Re: A dataset of every Reddit comment
#16BigQuery is the best interface for it. Can resolve queries on the entire dataset in less than a few seconds (however, you only get 1TB processing free per month. Since the full dataset is ~285GB, you only get 4 queries per month. Plan ahead on the May 2015 dataset, which is only 8GB.)
I can answer any other questions that people have.
Re: A dataset of every Reddit comment
#17If anyone is working with the data, we should see how many users actually hivemind comment on things vs OC
But by sheer coincidence, I have made a chart comparing average scores for OC submissions vs. non-OC submissions a few months ago: https://www.reddit.com/r/dataisbeautiful/comments/2rv76z/oc_...
Re: A dataset of every Reddit comment
#18Re: A dataset of every Reddit comment
#19Re: A dataset of every Reddit comment
#20This could be amazing as input to a question answering engine.
[2001 HAL voice]: I have a question about the penis mightier. Does it work?