Earlier quoted context omitted.
Does BOSS still exist? I was under the impression that it was defunct.
Yes. With Google no longer providing search result API (not even paid version, the last I checked) people are turning to BOSS/Bing/(anything else?)
Free 5 Billion Page Web Index Now Available from Common Crawl Foundation
11–20 of 42 posts
Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation
#12I'd love to see Gabriel weigh in on this. I wonder if Duck Duck Go will be able to take advantage of this resource?
Nova Spivack said that the crawls have been going for several years. There's a good chance that many of the pages in the archive are unacceptably outdated for indexing purposes.
But I am sure that it would be logic failure to conclude that it must be out of date simply because they've been indexing for several years. With that logic, Google would be further out of date, having indexed for over a decade.
Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation
#13Take crawl freshness. If I publish a new blog post, it gets crawled and added to the Google index in seconds. Other crawling efforts take weeks between refreshes.
Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation
#14I'll just leave this here: http://training.fogcreek.com
Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation
#15I'd love to see Gabriel weigh in on this. I wonder if Duck Duck Go will be able to take advantage of this resource?
Nova Spivack said that the crawls have been going for several years. There's a good chance that many of the pages in the archive are unacceptably outdated for indexing purposes.
Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation
#16"Well this has to be a first for a software company" I'll just leave this here: http://training.fogcreek.com
Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation
#17I'd love to see Gabriel weigh in on this. I wonder if Duck Duck Go will be able to take advantage of this resource?
Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation
#18One interesting discussion from here: http://www.commoncrawl.org/common-crawl-enters-a-new-phase/ It says the cost of running a hadoop job to scan all 5billon documents is in the order of $100. Does any one know how does this compare to let say Yahoo BOSS? Is it even comparable?
As far as comparisons to Yahoo BOSS are concerned, no, we are definitely not comparable to Yahoo BOSS or other such APIs that run on top of an already built (and properly ranked) inverted index of the web. At this stage we only produce bulk snapshots of what we crawl, and we are focusing our engineering resources on improving the frequency and coverage of crawl (the results of which will hopefully start to bear fruit in early 2012). Perhaps at some point in the near future, we can partner with the community to build a rudimentary full-text inverted index of the Web that we can make available in bulk via S3 as well.
Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation
#19I don't see any links to download their Hadoop classes..
Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation
#20I was hoping for Yahoo, Amazon, or Microsoft to throw a lot of resources at this about 5~8 years ago. Since then, Google kind of ran away with the game in crawling. They were far ahead of everyone else back then, but one could conceive of a rag-tag group of companies, institutions, and individuals pooling their resources and getting a crawl about 10% as good. These days, on the externally visible evidence they're pro…