I was in the process of reading this when I thought to check who this person is. Of course, by that time the site had failed, so I haven't read the whole thing yet. But, it seems to me that the author is falling in to a trap many an unwary data "scientist" falls by not understanding the discipline of Statistics. When one has the entire population data (i.e. a census), rather than a sample, there is no point in carryi…
Also, I am going to go out on a limb here and guess that R's `read.csv` doesn't do what one hopes it would when fed this CSV: 10,3,Brian,"You mean like the time you had tea with Mohammad, the prophet of the Muslim faith? Peter: Come on, Mohammad, let's get some tea. Mr. T: Try my ""Mr. T. ...tea."" " Well, it seems people are not understanding the problem with this line. Here is the screenshot of the original script:…
The data sources are CSVs in this repository: https://github.com/BobAdamsEE/SouthParkData/
Looks like all the data is preprocessed, with everyone mostly having only 1 line. (Actually, it appears the line you note in 10-3 is broken!) You can make an argument that the script isn't processed correctly, but that's beyond the scope of the analysis, although a note might be helpful.