Live data from Hacker News

Viewing profile — ccgreg

ccgreg

HN member
Joined
Thu, Nov 16, 2023, 9:21 PM UTC
HN karma
238
Public activity
130 items

About ccgreg

CTO at the Common Crawl Foundation

Recent public activity

  1. comment
    Comment #49219800

    > It's a shame that CCBot is caught in the cross fire, but that's life. We're used to it. Sadly.

  2. comment
    Comment #49219795

    The author blocked CCBot even though CCBot isn't part of the high traffic problem -- apparently he trusted Cloudflare labeling us as an "AI Bot".

  3. comment
    Comment #49187005

    That’s already a big business!

  4. comment
    Comment #49184503

    Thanks for the heads up -- this isn't popular yet, and it requires some work to avoid polluting things like the Internet Archive Wayback Machine.

  5. comment
    Comment #49172626

    The post says: > Because Common Crawl stores only the first 1 MB of each PDF That limit became 5 MB in March 2025.

  6. comment
    Comment #49139230

    Appreciate you double-checking.

  7. comment
    Comment #49138439

    That isn't true. You're welcome to peruse our index to prove or disprove your claim.

  8. comment
    Comment #48916196

    It's kind of interesting that you contradict much of what the article concludes, even though the article gives a lot of examples. Maybe your prediction will be true.

  9. comment
    Comment #48869598

    We do. First off we have a public parquet-format index of all of the urls we crawl every month. And then that also lives in a HDFS table that determines when we want to recrawl a p…

  10. comment
    Comment #48869583

    Appreciate your kind words! Many people have worked at Common Crawl over the years, and it's been a labor of love fueled by positive comments like yours and the large list of PhD t…

  11. comment
    Comment #48868920

    If you're referring to Common Crawl, which has existed since 2008, indeed your predictions are somewhat accurate. It's easy to opt out or limit what is collected. The crawling itse…

  12. comment
    Comment #48868904

    We aren't sure if that really made a significant difference in Common Crawl's data quality. It does hurt our dataset from a humanities point of view, alas.

  13. comment
    Comment #48868249

    Common Crawl's archive has metadata that says when each record (html file) was crawled.

  14. comment
    Comment #48868244

    Common Crawl's dataset was downloaded in full 100 times in 2025. We agree that it would be great if it was even more widely used.

  15. comment
    Comment #48866005

    A lot of websites want "bot defense" due to high volume scrapers, and that "bot defense" often also ends up blocking low-volume wget/curl and polite crawlers like Common Crawl's CC…

  16. comment
    Comment #48343285

    Good timing, I'm about to release that dataset.

  17. comment
    Comment #48317725

    Common Crawl is working hard to improve diversity in our crawl.

  18. comment
    Comment #47915461

    I don't know of anyone who uses Common Crawl as pre-training data without filtering it. We have an annotation system that lets people pick and choose which subsets they'd like to u…

  19. comment
    Comment #47907853

    Common Crawl is a sample of the web, so it's not that directly helpful for someone wanting to make a product price dataset.

  20. comment
    Comment #47907826

    I'm a life-long hacker, and my crawler crawls with consent.

  21. comment
    Comment #47795277

    The largest index we had was 4 billion, which is tiny. Our crawl frontier was much larger.

  22. comment
    Comment #47787571

    > and the data that I’ve experimented with from 2014 seemed high quality That's because it's from the blekko search engine.

  23. comment
    Comment #47715766

    That's already been happening for more than a year now.

  24. comment
    Comment #47598020

    Common Crawl is switching to reporting dataset sizes in nibbles. As an organisation dedicated to data preservation, we feel it would be remiss to allow this underrepresented unit t…

  25. story