Live data from Hacker News

Free 5 Billion Page Web Index Now Available from Common Crawl Foundation

readwriteweb.com

11–20 of 42 posts

Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation

#11
post #5
post #3

Earlier quoted context omitted.

Does BOSS still exist? I was under the impression that it was defunct.

Yes. With Google no longer providing search result API (not even paid version, the last I checked) people are turning to BOSS/Bing/(anything else?)

custom search API is the search result APi. The cse has a flag for searching the entire internet. http://www.google.com/support/customsearch/bin/answer.py?hl=...

Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation

#12
post #7

I'd love to see Gabriel weigh in on this. I wonder if Duck Duck Go will be able to take advantage of this resource?

Nova Spivack said that the crawls have been going for several years. There's a good chance that many of the pages in the archive are unacceptably outdated for indexing purposes.

I'm not sure whether major portions of their archive are unacceptably outdated.

But I am sure that it would be logic failure to conclude that it must be out of date simply because they've been indexing for several years. With that logic, Google would be further out of date, having indexed for over a decade.

Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation

#13
I was hoping for Yahoo, Amazon, or Microsoft to throw a lot of resources at this about 5~8 years ago. Since then, Google kind of ran away with the game in crawling. They were far ahead of everyone else back then, but one could conceive of a rag-tag group of companies, institutions, and individuals pooling their resources and getting a crawl about 10% as good. These days, on the externally visible evidence they're probably several orders of magnitude better than everybody else on the planet combined.

Take crawl freshness. If I publish a new blog post, it gets crawled and added to the Google index in seconds. Other crawling efforts take weeks between refreshes.

Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation

#15
post #7

I'd love to see Gabriel weigh in on this. I wonder if Duck Duck Go will be able to take advantage of this resource?

Nova Spivack said that the crawls have been going for several years. There's a good chance that many of the pages in the archive are unacceptably outdated for indexing purposes.

Hi. I work for commoncrawl. We are about to start an improved recrawl and will be doing this more frequently going forward. In the process we will also consolidate our data on S3 to keep it relevant. But, as with any crawl of the Internet, there is lot of noise in there. We spent most of 2011 tweaking the algorithms to improve the freshness and quality of the crawl, and hopefully this work starts to show results in 2012.

Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation

#17
post #7

I'd love to see Gabriel weigh in on this. I wonder if Duck Duck Go will be able to take advantage of this resource?

This was my first thought, too. I can't seem to hit the resource to check it out but even if the content is "stale" it can be used for a couple different reasons. Initial snapshot of pages, decent starting index on self crawling (instead of reliance on BOSS or Bing), content differentiation, nullifying search bias (if existent)... But I'm not really a search guy so I could be jaded on its importance.

Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation

#18
post #2

One interesting discussion from here: http://www.commoncrawl.org/common-crawl-enters-a-new-phase/ It says the cost of running a hadoop job to scan all 5billon documents is in the order of $100. Does any one know how does this compare to let say Yahoo BOSS? Is it even comparable?

Hi, I work at commoncrawl, so I will try to answer your question. We store our crawl data on S3 in the form of 100MB compressed archives and there are between 40,000 and 50,000 such files in commoncrawl’s bucket today. The key to scanning such a large set of files efficiently on EC2 is to have your each of your Mappers (assuming you are running Hadoop) open multiple S3 streams in parallel to maintain some desired level of throughput. For example, assuming that you can maintain on average a 1MByte/sec throughput per S3 stream, and you start 10 parallel streams per Mapper, you should be able to sustain a throughput 80 Mbits/sec or 10 MBytes/sec. If you were to run one Mapper per EC2 small instance, and start 100 such instances, this would yield and aggregated throughput of close to 3TB/hour. At that rate, you would need 16 hours to scan 50TB of data, or a total of 1600 machine hours at $.085 per hour, costing you somewhere in the neighborhood of $130.00. Of course, you would then need to add in the cost of running any subsequent aggregation / data consolidation jobs and the cost of storing your final data on S3. So, the $100.00 number is generally in the ballpark but final numbers may vary :-)

As far as comparisons to Yahoo BOSS are concerned, no, we are definitely not comparable to Yahoo BOSS or other such APIs that run on top of an already built (and properly ranked) inverted index of the web. At this stage we only produce bulk snapshots of what we crawl, and we are focusing our engineering resources on improving the frequency and coverage of crawl (the results of which will hopefully start to bear fruit in early 2012). Perhaps at some point in the near future, we can partner with the community to build a rudimentary full-text inverted index of the Web that we can make available in bulk via S3 as well.

Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation

#20
post #13

I was hoping for Yahoo, Amazon, or Microsoft to throw a lot of resources at this about 5~8 years ago. Since then, Google kind of ran away with the game in crawling. They were far ahead of everyone else back then, but one could conceive of a rag-tag group of companies, institutions, and individuals pooling their resources and getting a crawl about 10% as good. These days, on the externally visible evidence they're pro…

Hi I work at commoncrawl. We have spent our time (in 2011) improving our algorithms, and hopefully this effort will start to show real results (with respect to crawl frequency and relevancy) in 2012. But you are right, it is pretty unlikely that our crawl will be able to be fully competitive with the likes of Google etc., multi-billion dollar corporations who dedicate huge amounts of engineering and hardware resources to stay competitive in this field.
Post reply on HN