Live data from Hacker News

Classifying all of the pdfs on the internet

snats.xyz

11–20 of 117 posts

Re: Classifying all of the pdfs on the internet

#12
post #2

Interesting read, I did not know about Common Crawl. I feel like RTBF is kind of a lost battle these days with more and more crawlers for AI and whatnot. Once on the internet there is no way back, for better or for worse. This tangent aside, 8TB is really not a lot of data, it's just 8 consumer-grade 1TB hard drives. I find it hard to believe this is "the largest corpus of PDFs online", maybe the largest public one.…

Libgen size is ~33TB so, no, it's not "the largest corpus of PDFs online". (Although you could argue libgen is not really "public" in the legal sense of the word, lol). Disregarding that, the article is great! (edit: why would someone downvote this, HN is becoming quite hostile lately)

It's being down voted because your number is really off. Libgen's corpus is 100+ TB

Re: Classifying all of the pdfs on the internet

#13
post #2

Interesting read, I did not know about Common Crawl. I feel like RTBF is kind of a lost battle these days with more and more crawlers for AI and whatnot. Once on the internet there is no way back, for better or for worse. This tangent aside, 8TB is really not a lot of data, it's just 8 consumer-grade 1TB hard drives. I find it hard to believe this is "the largest corpus of PDFs online", maybe the largest public one.…

Libgen size is ~33TB so, no, it's not "the largest corpus of PDFs online". (Although you could argue libgen is not really "public" in the legal sense of the word, lol). Disregarding that, the article is great! (edit: why would someone downvote this, HN is becoming quite hostile lately)

I haven't downvoted you but it is presumably because of your hasty typing or lack of proofreading/research.

33TB (first google result from 5 years ago) not 33GB. Larger figures from more recently.

Re: Classifying all of the pdfs on the internet

#14

I have 20-40TB (pre-dedup) of PDFs - 8TB is a lot but not even close to the total number of PDFs available.

Just wondering what do you collect? Is it mainly mirroring things like libgen?

I have a decent collection of ebooks/pdfs/manga from reading. But I can’t imagine how large a 20TB library is.

Re: Classifying all of the pdfs on the internet

#15
post #2

Interesting read, I did not know about Common Crawl. I feel like RTBF is kind of a lost battle these days with more and more crawlers for AI and whatnot. Once on the internet there is no way back, for better or for worse. This tangent aside, 8TB is really not a lot of data, it's just 8 consumer-grade 1TB hard drives. I find it hard to believe this is "the largest corpus of PDFs online", maybe the largest public one.…

RTBF was a ludicrous concept before AI and these new crawlers.

Only EU bureaucracts would have the hubris to believe you could actually, comprehensively remove information from the Internet. Once something is spread, it is there, forever.

Re: Classifying all of the pdfs on the internet

#16
post #2

Interesting read, I did not know about Common Crawl. I feel like RTBF is kind of a lost battle these days with more and more crawlers for AI and whatnot. Once on the internet there is no way back, for better or for worse. This tangent aside, 8TB is really not a lot of data, it's just 8 consumer-grade 1TB hard drives. I find it hard to believe this is "the largest corpus of PDFs online", maybe the largest public one.…

> RTBF

Right to be forgotten, not the Belgian public service broadcaster (https://en.wikipedia.org/wiki/RTBF)?

Re: Classifying all of the pdfs on the internet

#18

I have 20-40TB (pre-dedup) of PDFs - 8TB is a lot but not even close to the total number of PDFs available.

Just wondering what do you collect? Is it mainly mirroring things like libgen? I have a decent collection of ebooks/pdfs/manga from reading. But I can’t imagine how large a 20TB library is.

No torrents at all in this data, all publicly available/open access. Mostly scientific pdfs, and a good portion of those are scans not just text. So the actual text amount is probably pretty low compared to the total. But still, a lot more than 8TB of raw data out there. I bet the total number of PDFs is close to a petabyte if not more.

Re: Classifying all of the pdfs on the internet

#19
post #9

Earlier quoted context omitted.

Libgen size is ~33TB so, no, it's not "the largest corpus of PDFs online". (Although you could argue libgen is not really "public" in the legal sense of the word, lol). Disregarding that, the article is great! (edit: why would someone downvote this, HN is becoming quite hostile lately)

8TB - ~8,000GB - is more than 33GB.

Whoops, typo!

But that's what the comments are for, not the downvotes.

Re: Classifying all of the pdfs on the internet

#20
post #2

Interesting read, I did not know about Common Crawl. I feel like RTBF is kind of a lost battle these days with more and more crawlers for AI and whatnot. Once on the internet there is no way back, for better or for worse. This tangent aside, 8TB is really not a lot of data, it's just 8 consumer-grade 1TB hard drives. I find it hard to believe this is "the largest corpus of PDFs online", maybe the largest public one.…

Libgen size is ~33TB so, no, it's not "the largest corpus of PDFs online". (Although you could argue libgen is not really "public" in the legal sense of the word, lol). Disregarding that, the article is great! (edit: why would someone downvote this, HN is becoming quite hostile lately)

I think Libgen is ~100TB, and the full Anna's Archive is near a PB.

They all probably contain lots of duplicates but...

https://annas-archive.se/datasets

Post reply on HN