Live data from Hacker News

Viewing profile — ahadrana

ahadrana

HN member
Joined
Thu, Aug 19, 2010, 10:32 PM UTC
HN karma
19
Public activity
7 items

About ahadrana

No profile information was provided.

Recent public activity

  1. comment
    Comment #3214570

    Hi, you can view our terms of use at http://www.commoncrawl.org/about/terms-of-use/full-terms-of-... . We adhere to the robots.txt standard, try to do all our crawling above board,…

  2. comment
    Comment #3211868

    We hear you. Could you define some criteria as to the type and size of sample data you would like to see? We are working on producing more targeted/limited collections, like perhap…

  3. comment
    Comment #3211084

    The pagerank and other metadata we compute is not part of the S3 corpus, but we do collect this information and probably will make it available in a separate S3 bucket in Hadoop Se…

  4. comment
    Comment #3210632

    Hi I work at commoncrawl. We have spent our time (in 2011) improving our algorithms, and hopefully this effort will start to show real results (with respect to crawl frequency and …

  5. comment
    Comment #3210601

    Sorry, our github repository had some accidental check-ins that we needed to remove. I will share the link to the code shortly.

  6. comment
    Comment #3210592

    Hi, I work at commoncrawl, so I will try to answer your question. We store our crawl data on S3 in the form of 100MB compressed archives and there are between 40,000 and 50,000 suc…

  7. comment
    Comment #3210497

    Hi. I work for commoncrawl. We are about to start an improved recrawl and will be doing this more frequently going forward. In the process we will also consolidate our data on S3 t…