A Look Inside Our 210TB 2012 Web Corpus
commoncrawl.org
A Look Inside Our 210TB 2012 Web Corpus
1–10 of 38 posts
Re: A Look Inside Our 210TB 2012 Web Corpus
#2Re: A Look Inside Our 210TB 2012 Web Corpus
#3Re: A Look Inside Our 210TB 2012 Web Corpus
#4What do you mean by "open"? Can the data be used for startups and other commercial purposes?
and you just need to comply with the Common Crawl TOU: http://commoncrawl.org/about/terms-of-use/
Re: A Look Inside Our 210TB 2012 Web Corpus
#5What do you mean by "open"? Can the data be used for startups and other commercial purposes?
Re: A Look Inside Our 210TB 2012 Web Corpus
#6Re: A Look Inside Our 210TB 2012 Web Corpus
#7What do you mean by "open"? Can the data be used for startups and other commercial purposes?
Actually, tomorrow a video on a startup that uses Common Crawl data is getting posted.
Re: A Look Inside Our 210TB 2012 Web Corpus
#8Re: A Look Inside Our 210TB 2012 Web Corpus
#9How does one get set up to access the s3:// links their blog posts reference? I do realize these point to Amazon S3 buckets, but how to get at them?
From there you can grab the S3 command line tools (http://s3tools.org/s3cmd) or load it up from hadoop or through one of the various open source libraries (boto for instance).
Re: A Look Inside Our 210TB 2012 Web Corpus
#10Is there something, other than funding, preventing a more regular, open-sourced crawl of the web?