One interesting discussion from here: http://www.commoncrawl.org/common-crawl-enters-a-new-phase/ It says the cost of running a hadoop job to scan all 5billon documents is in the order of $100. Does any one know how does this compare to let say Yahoo BOSS? Is it even comparable?
Hi, I work at commoncrawl, so I will try to answer your question. We store our crawl data on S3 in the form of 100MB compressed archives and there are between 40,000 and 50,000 such files in commoncrawl’s bucket today. The key to scanning such a large set of files efficiently on EC2 is to have your each of your Mappers (assuming you are running Hadoop) open multiple S3 streams in parallel to maintain some desired lev…
Free 5 Billion Page Web Index Now Available from Common Crawl Foundation
21–30 of 42 posts
Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation
#22I wonder if crooks will try to exploit this crawl. As a person who has an index of the web like this it has been interesting to see what they look for. SSN's and credit card numbers are common, as are sites running older versions of PHP software or exploitable shopping carts.
It makes it very easy for people to steal vast amounts of your content and republish it on their own sites, with ads all around it. Many content sites have protections in place to recognize bots by their behavior or use "honeypots" to tell bots apart from human visitors and thus avoid large scale content theft.
Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation
#23I was hoping for Yahoo, Amazon, or Microsoft to throw a lot of resources at this about 5~8 years ago. Since then, Google kind of ran away with the game in crawling. They were far ahead of everyone else back then, but one could conceive of a rag-tag group of companies, institutions, and individuals pooling their resources and getting a crawl about 10% as good. These days, on the externally visible evidence they're pro…
In the 2004 timeframe, Yahoo was crawling about the same number of pages as Google. (More some months.)
> If I publish a new blog post, it gets crawled and added to the Google index in seconds. Other crawling efforts take weeks between refreshes.
Time from crawl to appearing in search results is a different issue.
Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation
#24Earlier quoted context omitted.
Hi, I work at commoncrawl, so I will try to answer your question. We store our crawl data on S3 in the form of 100MB compressed archives and there are between 40,000 and 50,000 such files in commoncrawl’s bucket today. The key to scanning such a large set of files efficiently on EC2 is to have your each of your Mappers (assuming you are running Hadoop) open multiple S3 streams in parallel to maintain some desired lev…
Hey ahadrana, I haven't found anything about the page ranks on the website, are they included? Do you know if it is possible to go only trough the metadata of the crawl, say to get the page ranks for a list of pages or do you have to go through the full crawl?
Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation
#25I'd love to see Gabriel weigh in on this. I wonder if Duck Duck Go will be able to take advantage of this resource?
Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation
#26The crawl file presumably contains the contents of websites and so the owners of those websites could assert that Common Crawl Foundation is distributing their work without permission or license.
There are all sorts of republishing/splog 'opportunities' with this crawl data that goes beyond the original expected use.
Surprisingly, I couldn't see anything about this covered in the FAQs
Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation
#27Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation
#28I was hoping for Yahoo, Amazon, or Microsoft to throw a lot of resources at this about 5~8 years ago. Since then, Google kind of ran away with the game in crawling. They were far ahead of everyone else back then, but one could conceive of a rag-tag group of companies, institutions, and individuals pooling their resources and getting a crawl about 10% as good. These days, on the externally visible evidence they're pro…
Hi I work at commoncrawl. We have spent our time (in 2011) improving our algorithms, and hopefully this effort will start to show real results (with respect to crawl frequency and relevancy) in 2012. But you are right, it is pretty unlikely that our crawl will be able to be fully competitive with the likes of Google etc., multi-billion dollar corporations who dedicate huge amounts of engineering and hardware resource…
Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation
#29I think all projects should have sample datasets. It simplifies a lot of things, and in this case stops hundreds of geeks burning through bandwidth before they realize they don't have a clue what they are going to do with the data.
Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation
#30Is there a sample dataset? I think all projects should have sample datasets. It simplifies a lot of things, and in this case stops hundreds of geeks burning through bandwidth before they realize they don't have a clue what they are going to do with the data.