CommonCrawl: an open repository of web crawl data that is universally accessible
1–10 of 10 posts
Re: CommonCrawl: an open repository of web crawl data that is universally accessible
#2Just a pointer, the code for CommonCrawl Project is available on Github https://github.com/commoncrawl/commoncrawl
Re: CommonCrawl: an open repository of web crawl data that is universally accessible
#3This looks really nice!
Re: CommonCrawl: an open repository of web crawl data that is universally accessible
#4I hear a lot of people are crunching on CommonCrawl's data. It'll be interesting the type of stuff people come up with!
Re: CommonCrawl: an open repository of web crawl data that is universally accessible
#5If you into said things then maybe http://yacy.net/ (p2p crawler and search) will be useful to you as well.
Re: CommonCrawl: an open repository of web crawl data that is universally accessible
#6thread on HN from when common crawl was announced, interesting info there:
http://news.ycombinator.com/item?id=3209690
Re: CommonCrawl: an open repository of web crawl data that is universally accessible
#7It would be great to hear more about the tools they are using to crawl and potentially open it up to more people who want to contribute computing resources.
Re: CommonCrawl: an open repository of web crawl data that is universally accessible
#8[deleted]
Re: CommonCrawl: an open repository of web crawl data that is universally accessible
#9The latest data available is from 2010-09-25, which seems to be too old to be useful for most things.
Re: CommonCrawl: an open repository of web crawl data that is universally accessible
#10This one may be also interesting for open data devs:
http://scraperwiki.com/