Processing 40 TB of code from ~10M projects with a server and Go for $100 (2019)
11–16 of 16 posts
Re: Processing 40 TB of code from ~10M projects with a server and Go for $100 (2019)
#12Re: Processing 40 TB of code from ~10M projects with a server and Go for $100 (2019)
#13Average number of Makefiles is 1518.1769808494607? That can’t be right?
Re: Processing 40 TB of code from ~10M projects with a server and Go for $100 (2019)
#14>>> If someone wants to host the raw files to allow others to download it let me know. It is a 83 GB tar.gz file which uncompressed is just over 1 TB in size. Anyone knows if it is possible to download similar data set for youtube and reddit? I have ideas for search engine based on it, but I don't want to write/maintain scraper scripts.
Re: Processing 40 TB of code from ~10M projects with a server and Go for $100 (2019)
#15Average number of Makefiles is 1518.1769808494607? That can’t be right?
Re: Processing 40 TB of code from ~10M projects with a server and Go for $100 (2019)
#16Average number of Makefiles is 1518.1769808494607? That can’t be right?
The ”average person eats 3 spiders a year” factoid is actually just statistical error. The average person eats 0 spiders per year. Spiders Georg, who lives in a cave & eats over 10,000 each day, is an outlier and should not have been counted”
Also, according to the page, only 59 million of all files are named “Makefile”.
I suspect that the file language recognition has major false positives for the Makefile language.