102TB of New Crawl Data Available
commoncrawl.org
102TB of New Crawl Data Available
1–10 of 39 posts
Re: 102TB of New Crawl Data Available
#2I really think a subset like this will increase the value as it would allow people writing search engines (for fun or profit) to suck a copy down locally and work away. Its something I would like to do for sure.
Re: 102TB of New Crawl Data Available
#3I love common crawl, but as I commented before I still want to see a subset available for download, something like the top million sites or something like that. Certainly a few steps of data, say 50GB 100GB and 200GB. I really think a subset like this will increase the value as it would allow people writing search engines (for fun or profit) to suck a copy down locally and work away. Its something I would like to do…
Re: 102TB of New Crawl Data Available
#4I love common crawl, but as I commented before I still want to see a subset available for download, something like the top million sites or something like that. Certainly a few steps of data, say 50GB 100GB and 200GB. I really think a subset like this will increase the value as it would allow people writing search engines (for fun or profit) to suck a copy down locally and work away. Its something I would like to do…
There will be news about a subset sometime next month!
Re: 102TB of New Crawl Data Available
#5Re: 102TB of New Crawl Data Available
#6Re: 102TB of New Crawl Data Available
#7I love common crawl, but as I commented before I still want to see a subset available for download, something like the top million sites or something like that. Certainly a few steps of data, say 50GB 100GB and 200GB. I really think a subset like this will increase the value as it would allow people writing search engines (for fun or profit) to suck a copy down locally and work away. Its something I would like to do…
There will be news about a subset sometime next month!
Re: 102TB of New Crawl Data Available
#8Very cool...though I have to say, CC is a constant reminder that whatever you put on the Internet will basically remain in the public eye for the perpetuity of electronic communication. There exists ways to remove your (owned) content from archive.org and Google...but once some other independent scraper catches it, you can't really do much about it
Re: 102TB of New Crawl Data Available
#9It looks like they've fixed the first problem by switching to gzipped WARC files, but I can't find any information about whether or not they're still truncating documents in the archive. I guess I'll have to give it another look and see...
Re: 102TB of New Crawl Data Available
#10Earlier quoted context omitted.
There will be news about a subset sometime next month!
would love to have even smaller subsets (like 5gb) that students can casually play around with too to practice and learn tools and algos :) (if it's not too much trouble!)