Classifying all of the pdfs on the internet
1–10 of 117 posts
Re: Classifying all of the pdfs on the internet
#2Re: Classifying all of the pdfs on the internet
#3Interesting read, I did not know about Common Crawl. I feel like RTBF is kind of a lost battle these days with more and more crawlers for AI and whatnot. Once on the internet there is no way back, for better or for worse. This tangent aside, 8TB is really not a lot of data, it's just 8 consumer-grade 1TB hard drives. I find it hard to believe this is "the largest corpus of PDFs online", maybe the largest public one.…
If you like, you could say, PDF are information dense, but data sparse. After all it is mostly white space ;)
Re: Classifying all of the pdfs on the internet
#4Interesting read, I did not know about Common Crawl. I feel like RTBF is kind of a lost battle these days with more and more crawlers for AI and whatnot. Once on the internet there is no way back, for better or for worse. This tangent aside, 8TB is really not a lot of data, it's just 8 consumer-grade 1TB hard drives. I find it hard to believe this is "the largest corpus of PDFs online", maybe the largest public one.…
(Although you could argue libgen is not really "public" in the legal sense of the word, lol).
Disregarding that, the article is great!
(edit: why would someone downvote this, HN is becoming quite hostile lately)
Re: Classifying all of the pdfs on the internet
#5Interesting read, I did not know about Common Crawl. I feel like RTBF is kind of a lost battle these days with more and more crawlers for AI and whatnot. Once on the internet there is no way back, for better or for worse. This tangent aside, 8TB is really not a lot of data, it's just 8 consumer-grade 1TB hard drives. I find it hard to believe this is "the largest corpus of PDFs online", maybe the largest public one.…
[0] https://link.springer.com/article/10.1007/s11192-015-1614-6
Re: Classifying all of the pdfs on the internet
#6Interesting read, I did not know about Common Crawl. I feel like RTBF is kind of a lost battle these days with more and more crawlers for AI and whatnot. Once on the internet there is no way back, for better or for worse. This tangent aside, 8TB is really not a lot of data, it's just 8 consumer-grade 1TB hard drives. I find it hard to believe this is "the largest corpus of PDFs online", maybe the largest public one.…
Re: Classifying all of the pdfs on the internet
#7Interesting read, I did not know about Common Crawl. I feel like RTBF is kind of a lost battle these days with more and more crawlers for AI and whatnot. Once on the internet there is no way back, for better or for worse. This tangent aside, 8TB is really not a lot of data, it's just 8 consumer-grade 1TB hard drives. I find it hard to believe this is "the largest corpus of PDFs online", maybe the largest public one.…
Doesn't sound like a lot, but where I am now we routinely work on very large infrastructure projects and the plans, documents and stuff mostly come as PDF. We are talking of thousands of documents, often with thousands of pages, per project and even very big projects almost never break 20 GB. If you like, you could say, PDF are information dense, but data sparse. After all it is mostly white space ;)
Re: Classifying all of the pdfs on the internet
#8Interesting read, I did not know about Common Crawl. I feel like RTBF is kind of a lost battle these days with more and more crawlers for AI and whatnot. Once on the internet there is no way back, for better or for worse. This tangent aside, 8TB is really not a lot of data, it's just 8 consumer-grade 1TB hard drives. I find it hard to believe this is "the largest corpus of PDFs online", maybe the largest public one.…
For those of us who aren't familiar with this random acronym, I think RTBF = right to be forgotten.
Re: Classifying all of the pdfs on the internet
#9Interesting read, I did not know about Common Crawl. I feel like RTBF is kind of a lost battle these days with more and more crawlers for AI and whatnot. Once on the internet there is no way back, for better or for worse. This tangent aside, 8TB is really not a lot of data, it's just 8 consumer-grade 1TB hard drives. I find it hard to believe this is "the largest corpus of PDFs online", maybe the largest public one.…
Libgen size is ~33TB so, no, it's not "the largest corpus of PDFs online". (Although you could argue libgen is not really "public" in the legal sense of the word, lol). Disregarding that, the article is great! (edit: why would someone downvote this, HN is becoming quite hostile lately)
Re: Classifying all of the pdfs on the internet
#10Interesting read, I did not know about Common Crawl. I feel like RTBF is kind of a lost battle these days with more and more crawlers for AI and whatnot. Once on the internet there is no way back, for better or for worse. This tangent aside, 8TB is really not a lot of data, it's just 8 consumer-grade 1TB hard drives. I find it hard to believe this is "the largest corpus of PDFs online", maybe the largest public one.…
Is it possible that the 8 TB is just the extracted text?