Live data from Hacker News

Classifying all of the pdfs on the internet

snats.xyz

31–40 of 117 posts

Re: Classifying all of the pdfs on the internet

#31
post #15
post #2

Interesting read, I did not know about Common Crawl. I feel like RTBF is kind of a lost battle these days with more and more crawlers for AI and whatnot. Once on the internet there is no way back, for better or for worse. This tangent aside, 8TB is really not a lot of data, it's just 8 consumer-grade 1TB hard drives. I find it hard to believe this is "the largest corpus of PDFs online", maybe the largest public one.…

RTBF was a ludicrous concept before AI and these new crawlers. Only EU bureaucracts would have the hubris to believe you could actually, comprehensively remove information from the Internet. Once something is spread, it is there, forever.

RTBF isn't about having your information wiped from the internet. Its a safe assumption any public information about you is completely out of your control as soon as its public.

RTBF is about getting companies to get rid of any trace of you so they cannot use that data, not removing all traces about you across the internet.

Re: Classifying all of the pdfs on the internet

#32

I have 20-40TB (pre-dedup) of PDFs - 8TB is a lot but not even close to the total number of PDFs available.

Care to make it publicly available? Or is that not permitted on your dataset? Certainly, there’s a lot more PDFs out there than 8TB. I bet there’s a lot of redundancy in yours, but doesn’t dedup well because of all the images.

Re: Classifying all of the pdfs on the internet

#33
post #26
post #15

Earlier quoted context omitted.

RTBF was a ludicrous concept before AI and these new crawlers. Only EU bureaucracts would have the hubris to believe you could actually, comprehensively remove information from the Internet. Once something is spread, it is there, forever.

Correct me if I'm wrong, but I always took RTBF to mean you have the right to be forgotten by any specific service provider: that you can request they delete the data they have that relates to you, and that they forward the request to any subprocessors. That's fairly reasonable and doable, it is enforced by GDPR and a number of other wide-reaching laws already, and it is a relatively common practice nowadays to allow…

There is a whole business sector for ”Online reputation fixers”

https://www.mycleanslate.co.uk/

What they usually do

- Spam Google with the name to bury content

- Send legal threads and use GDPR

They have legit use cases, but are often used by convicted or shady businessmen, politicians, and scammers to hide their earlier misdeeds.

Re: Classifying all of the pdfs on the internet

#34

Earlier quoted context omitted.

I haven't downvoted you but it is presumably because of your hasty typing or lack of proofreading/research. 33TB (first google result from 5 years ago) not 33GB. Larger figures from more recently.

>hasty typing or lack of proofreading/research This is exactly what I meant with "HN is becoming quite hostile" * I brought up something I looked up to support GP's argument. * The argument is correct. * I do it in good faith. * G is literally next to T. * I even praise the article, while at it. "Oh, but you made a typo!". Good luck, guys. I'm out. PS. I will give my whole 7 figure net worth, no questions asked, tran…

> I will give my whole 7 figure net worth

You sound deeply unpleasant to talk to.

Imaginary internet points are just that.

Re: Classifying all of the pdfs on the internet

#35
post #2

Interesting read, I did not know about Common Crawl. I feel like RTBF is kind of a lost battle these days with more and more crawlers for AI and whatnot. Once on the internet there is no way back, for better or for worse. This tangent aside, 8TB is really not a lot of data, it's just 8 consumer-grade 1TB hard drives. I find it hard to believe this is "the largest corpus of PDFs online", maybe the largest public one.…

Libgen size is ~33TB so, no, it's not "the largest corpus of PDFs online". (Although you could argue libgen is not really "public" in the legal sense of the word, lol). Disregarding that, the article is great! (edit: why would someone downvote this, HN is becoming quite hostile lately)

(edit: why would someone downvote this, HN is becoming quite hostile lately)

Also, there are browser extensions that will automatically downvote and/or hide HN comments that use words like "lol," or start with "So..." or include any of a number of words that the user considers indicative of low-grade content.

Re: Classifying all of the pdfs on the internet

#36

I have 20-40TB (pre-dedup) of PDFs - 8TB is a lot but not even close to the total number of PDFs available.

Just wondering what do you collect? Is it mainly mirroring things like libgen? I have a decent collection of ebooks/pdfs/manga from reading. But I can’t imagine how large a 20TB library is.

Just wondering what do you collect?

I can't speak for the OP, but you can buy optical media of old out-of-print magazines scanned as PDFs.

I bought the entirety of Desert Magazine from 1937-1985. It arrived on something like 15 CD-ROMS.

I drag-and-dropped the entire collection into iBooks, and read them when I'm on the train.

(Yes, they're probably on archive.org for free, but this is far easier and more convenient, and I prefer to support publishers rather than undermine their efforts.)

Re: Classifying all of the pdfs on the internet

#37
This is a really cool idea, thanks for sharing. I don't have that much free time these days, but I was thinking of trying a similar-but-different project not too long ago.

I wanted to make a bit of an open source tool to pull down useful time series data for the social sciences (e.g. time series of social media comments about grocery prices). Seems like LLMs have unlocked all kinds of new research angles that people aren't using yet.

I may steal some of your good ideas if I ever get to work on that side project :)

Re: Classifying all of the pdfs on the internet

#39
post #15
post #2

Interesting read, I did not know about Common Crawl. I feel like RTBF is kind of a lost battle these days with more and more crawlers for AI and whatnot. Once on the internet there is no way back, for better or for worse. This tangent aside, 8TB is really not a lot of data, it's just 8 consumer-grade 1TB hard drives. I find it hard to believe this is "the largest corpus of PDFs online", maybe the largest public one.…

RTBF was a ludicrous concept before AI and these new crawlers. Only EU bureaucracts would have the hubris to believe you could actually, comprehensively remove information from the Internet. Once something is spread, it is there, forever.

> Once something is spread, it is there, forever.

Really depends on the content. Tons of websites are going down everyday, link rot is a real thing. Internet archive or people don't save nearly everything.

Something I should do more often is saving mhtml copies of webpages I find interesting.

Re: Classifying all of the pdfs on the internet

#40
post #2

Interesting read, I did not know about Common Crawl. I feel like RTBF is kind of a lost battle these days with more and more crawlers for AI and whatnot. Once on the internet there is no way back, for better or for worse. This tangent aside, 8TB is really not a lot of data, it's just 8 consumer-grade 1TB hard drives. I find it hard to believe this is "the largest corpus of PDFs online", maybe the largest public one.…

Tangentially related, I was once handed a single PDF between 2 and 5 GBs in size and asked to run inference on it. This was the result of a miscommunication with the data provider, but I think it's funny and almost impressive that this file even exists.
Post reply on HN