Live data from Hacker News

Free 5 Billion Page Web Index Now Available from Common Crawl Foundation

readwriteweb.com

21–30 of 42 posts

Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation

#21
post #2

One interesting discussion from here: http://www.commoncrawl.org/common-crawl-enters-a-new-phase/ It says the cost of running a hadoop job to scan all 5billon documents is in the order of $100. Does any one know how does this compare to let say Yahoo BOSS? Is it even comparable?

Hi, I work at commoncrawl, so I will try to answer your question. We store our crawl data on S3 in the form of 100MB compressed archives and there are between 40,000 and 50,000 such files in commoncrawl’s bucket today. The key to scanning such a large set of files efficiently on EC2 is to have your each of your Mappers (assuming you are running Hadoop) open multiple S3 streams in parallel to maintain some desired lev…

Hey ahadrana, I haven't found anything about the page ranks on the website, are they included? Do you know if it is possible to go only trough the metadata of the crawl, say to get the page ranks for a list of pages or do you have to go through the full crawl?

Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation

#22
post #8
post #4

I wonder if crooks will try to exploit this crawl. As a person who has an index of the web like this it has been interesting to see what they look for. SSN's and credit card numbers are common, as are sites running older versions of PHP software or exploitable shopping carts.

It makes it very easy for people to steal vast amounts of your content and republish it on their own sites, with ads all around it. Many content sites have protections in place to recognize bots by their behavior or use "honeypots" to tell bots apart from human visitors and thus avoid large scale content theft.

Presumably those protections would prevent this bot from collecting data as well?

Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation

#23
post #13

I was hoping for Yahoo, Amazon, or Microsoft to throw a lot of resources at this about 5~8 years ago. Since then, Google kind of ran away with the game in crawling. They were far ahead of everyone else back then, but one could conceive of a rag-tag group of companies, institutions, and individuals pooling their resources and getting a crawl about 10% as good. These days, on the externally visible evidence they're pro…

> I was hoping for Yahoo, Amazon, or Microsoft to throw a lot of resources at this about 5~8 years ago. Since then, Google kind of ran away with the game in crawling.

In the 2004 timeframe, Yahoo was crawling about the same number of pages as Google. (More some months.)

> If I publish a new blog post, it gets crawled and added to the Google index in seconds. Other crawling efforts take weeks between refreshes.

Time from crawl to appearing in search results is a different issue.

Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation

#24
post #21

Earlier quoted context omitted.

Hi, I work at commoncrawl, so I will try to answer your question. We store our crawl data on S3 in the form of 100MB compressed archives and there are between 40,000 and 50,000 such files in commoncrawl’s bucket today. The key to scanning such a large set of files efficiently on EC2 is to have your each of your Mappers (assuming you are running Hadoop) open multiple S3 streams in parallel to maintain some desired lev…

Hey ahadrana, I haven't found anything about the page ranks on the website, are they included? Do you know if it is possible to go only trough the metadata of the crawl, say to get the page ranks for a list of pages or do you have to go through the full crawl?

The pagerank and other metadata we compute is not part of the S3 corpus, but we do collect this information and probably will make it available in a separate S3 bucket in Hadoop SequenceFiles format. Be aware that our pagerank will probably not have a high degree of correlation to Google's pagerank number, since their pagerank calculation is going to be a lot more sophisticated than our version.

Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation

#25
post #7

I'd love to see Gabriel weigh in on this. I wonder if Duck Duck Go will be able to take advantage of this resource?

I love stuff like this effort. The more open data sources, the better for everyone. I'm sure we (DuckDuckGo) will find a way to make use of it :)

Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation

#26
Although I'm personally all for open distribution of crawl data like this and all of my personal websites are CC-licensed, isn't there something to be said for the copyright status of the pages in the crawl file?

The crawl file presumably contains the contents of websites and so the owners of those websites could assert that Common Crawl Foundation is distributing their work without permission or license.

There are all sorts of republishing/splog 'opportunities' with this crawl data that goes beyond the original expected use.

Surprisingly, I couldn't see anything about this covered in the FAQs

Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation

#28
post #13

I was hoping for Yahoo, Amazon, or Microsoft to throw a lot of resources at this about 5~8 years ago. Since then, Google kind of ran away with the game in crawling. They were far ahead of everyone else back then, but one could conceive of a rag-tag group of companies, institutions, and individuals pooling their resources and getting a crawl about 10% as good. These days, on the externally visible evidence they're pro…

Hi I work at commoncrawl. We have spent our time (in 2011) improving our algorithms, and hopefully this effort will start to show real results (with respect to crawl frequency and relevancy) in 2012. But you are right, it is pretty unlikely that our crawl will be able to be fully competitive with the likes of Google etc., multi-billion dollar corporations who dedicate huge amounts of engineering and hardware resource…

It is not "Google etc., multi-billion dollar corporations" it is just Google.

Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation

#29
Is there a sample dataset?

I think all projects should have sample datasets. It simplifies a lot of things, and in this case stops hundreds of geeks burning through bandwidth before they realize they don't have a clue what they are going to do with the data.

Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation

#30
post #29

Is there a sample dataset? I think all projects should have sample datasets. It simplifies a lot of things, and in this case stops hundreds of geeks burning through bandwidth before they realize they don't have a clue what they are going to do with the data.

We hear you. Could you define some criteria as to the type and size of sample data you would like to see? We are working on producing more targeted/limited collections, like perhaps all most recently published blog posts etc.
Post reply on HN