PDF Trends 2026 Q2: Analysis of 20.6M PDFs from Common Crawl
1–9 of 9 posts
Re: PDF Trends 2026 Q2: Analysis of 20.6M PDFs from Common Crawl
#2The final dataset contains:
20,578,394 PDF documents approximately 38 TB of source data
Re: PDF Trends 2026 Q2: Analysis of 20.6M PDFs from Common Crawl
#3Re: PDF Trends 2026 Q2: Analysis of 20.6M PDFs from Common Crawl
#4What methodology did you use?
basic metadata: page count file size creation and modification dates PDF version (including Version entry in the document catalog) producer and creator encryption information and permissions annotations presence of interactive forms presence of optional content layers presence of digital signatures image only (scanned) pages Tagged PDF information: stats on the use of structure element types logical structure tree validation against ISO 32005
Re: PDF Trends 2026 Q2: Analysis of 20.6M PDFs from Common Crawl
#5Most PDFs remain relatively small, short documents. Encryption is uncommon and generally does not prevent document access. Accessibility text extraction is enabled in most encrypted documents, although a significant minority still disables it. PDF 1.7 continues to dominate document production. Proprietary extensions remain common, particularly in annotation workflows. Link annotations dominate all other annotation types combined.
Re: PDF Trends 2026 Q2: Analysis of 20.6M PDFs from Common Crawl
#6This first part of the June 2026 Common Crawl PDFs analysis reveals several long-term characteristics of PDF usage on the public web: Most PDFs remain relatively small, short documents. Encryption is uncommon and generally does not prevent document access. Accessibility text extraction is enabled in most encrypted documents, although a significant minority still disables it. PDF 1.7 continues to dominate document pro…
> Because Common Crawl stores only the first 1 MB of each PDF
That limit became 5 MB in March 2025.