Viewing profile — ccgreg
ccgreg
HN member- Joined
- Thu, Nov 16, 2023, 9:21 PM UTC
- HN karma
- 238
- Public activity
- 130 items
- HN profile
- View on Hacker News ↗
About ccgreg
Recent public activity
-
comment
Comment #49219800
> It's a shame that CCBot is caught in the cross fire, but that's life. We're used to it. Sadly.
-
comment
Comment #49219795
The author blocked CCBot even though CCBot isn't part of the high traffic problem -- apparently he trusted Cloudflare labeling us as an "AI Bot".
-
comment
Comment #49187005
That’s already a big business!
-
comment
Comment #49184503
Thanks for the heads up -- this isn't popular yet, and it requires some work to avoid polluting things like the Internet Archive Wayback Machine.
-
comment
Comment #49172626
The post says: > Because Common Crawl stores only the first 1 MB of each PDF That limit became 5 MB in March 2025.
-
comment
Comment #49139230
Appreciate you double-checking.
-
comment
Comment #49138439
That isn't true. You're welcome to peruse our index to prove or disprove your claim.
-
comment
Comment #48916196
It's kind of interesting that you contradict much of what the article concludes, even though the article gives a lot of examples. Maybe your prediction will be true.
-
comment
Comment #48869598
We do. First off we have a public parquet-format index of all of the urls we crawl every month. And then that also lives in a HDFS table that determines when we want to recrawl a p…
-
comment
Comment #48869583
Appreciate your kind words! Many people have worked at Common Crawl over the years, and it's been a labor of love fueled by positive comments like yours and the large list of PhD t…
-
comment
Comment #48868920
If you're referring to Common Crawl, which has existed since 2008, indeed your predictions are somewhat accurate. It's easy to opt out or limit what is collected. The crawling itse…
-
comment
Comment #48868904
We aren't sure if that really made a significant difference in Common Crawl's data quality. It does hurt our dataset from a humanities point of view, alas.
-
comment
Comment #48868249
Common Crawl's archive has metadata that says when each record (html file) was crawled.
-
comment
Comment #48868244
Common Crawl's dataset was downloaded in full 100 times in 2025. We agree that it would be great if it was even more widely used.
-
comment
Comment #48866005
A lot of websites want "bot defense" due to high volume scrapers, and that "bot defense" often also ends up blocking low-volume wget/curl and polite crawlers like Common Crawl's CC…
-
comment
Comment #48343285
Good timing, I'm about to release that dataset.
-
comment
Comment #48317725
Common Crawl is working hard to improve diversity in our crawl.
-
comment
Comment #47915461
I don't know of anyone who uses Common Crawl as pre-training data without filtering it. We have an annotation system that lets people pick and choose which subsets they'd like to u…
-
comment
Comment #47907853
Common Crawl is a sample of the web, so it's not that directly helpful for someone wanting to make a product price dataset.
-
comment
Comment #47907826
I'm a life-long hacker, and my crawler crawls with consent.
-
comment
Comment #47795277
The largest index we had was 4 billion, which is tiny. Our crawl frontier was much larger.
-
comment
Comment #47787571
> and the data that I’ve experimented with from 2014 seemed high quality That's because it's from the blekko search engine.
-
comment
Comment #47715766
That's already been happening for more than a year now.
-
comment
Comment #47598020
Common Crawl is switching to reporting dataset sizes in nibbles. As an organisation dedicated to data preservation, we feel it would be remiss to allow this underrepresented unit t…
- story