Viewing profile — ahadrana
ahadrana
HN member- Joined
- Thu, Aug 19, 2010, 10:32 PM UTC
- HN karma
- 19
- Public activity
- 7 items
- HN profile
- View on Hacker News ↗
About ahadrana
No profile information was provided.
Recent public activity
-
comment
Comment #3214570
Hi, you can view our terms of use at http://www.commoncrawl.org/about/terms-of-use/full-terms-of-... . We adhere to the robots.txt standard, try to do all our crawling above board,…
-
comment
Comment #3211868
We hear you. Could you define some criteria as to the type and size of sample data you would like to see? We are working on producing more targeted/limited collections, like perhap…
-
comment
Comment #3211084
The pagerank and other metadata we compute is not part of the S3 corpus, but we do collect this information and probably will make it available in a separate S3 bucket in Hadoop Se…
-
comment
Comment #3210632
Hi I work at commoncrawl. We have spent our time (in 2011) improving our algorithms, and hopefully this effort will start to show real results (with respect to crawl frequency and …
-
comment
Comment #3210601
Sorry, our github repository had some accidental check-ins that we needed to remove. I will share the link to the code shortly.
-
comment
Comment #3210592
Hi, I work at commoncrawl, so I will try to answer your question. We store our crawl data on S3 in the form of 100MB compressed archives and there are between 40,000 and 50,000 suc…
-
comment
Comment #3210497
Hi. I work for commoncrawl. We are about to start an improved recrawl and will be doing this more frequently going forward. In the process we will also consolidate our data on S3 t…