Live data from Hacker News

Classifying all of the pdfs on the internet

snats.xyz

61–70 of 117 posts

Re: Classifying all of the pdfs on the internet

#61
post #16
post #2

Interesting read, I did not know about Common Crawl. I feel like RTBF is kind of a lost battle these days with more and more crawlers for AI and whatnot. Once on the internet there is no way back, for better or for worse. This tangent aside, 8TB is really not a lot of data, it's just 8 consumer-grade 1TB hard drives. I find it hard to believe this is "the largest corpus of PDFs online", maybe the largest public one.…

> RTBF Right to be forgotten, not the Belgian public service broadcaster ( https://en.wikipedia.org/wiki/RTBF )?

Living in Belgium, I first thought that it was about the TV/radio service. Never saw the acronym R.T.B.F.

Re: Classifying all of the pdfs on the internet

#62
post #5
post #2

Interesting read, I did not know about Common Crawl. I feel like RTBF is kind of a lost battle these days with more and more crawlers for AI and whatnot. Once on the internet there is no way back, for better or for worse. This tangent aside, 8TB is really not a lot of data, it's just 8 consumer-grade 1TB hard drives. I find it hard to believe this is "the largest corpus of PDFs online", maybe the largest public one.…

Yeah 8TB is really tiny. Google scholar was estimated to index 160.000.000 pdfs in 2015.[0] If we assume that a third of those are not behind paywalls, and average pdf size is 1mb, its ends up as something above 50TB of documents. Almost ten years later the number of available pdfs of just scholarly communication should be substantially higher. [0] https://link.springer.com/article/10.1007/s11192-015-1614-6

Anna's archive has some 300M pdfs.

Re: Classifying all of the pdfs on the internet

#63
Did some similar work with similar visualizations ~2009, on ~5.7M research articles (PDFs, private corpus) from scientific publishers Elsevier, Springer:

Newton, G., A. Callahan & M. Dumontier. 2009. Semantic Journal Mapping for Search Visualization in a Large Scale Article Digital Library. Second Workshop on Very Large Digital Libraries at the European Conference on Digital Libraries (ECDL) 2009. https://lekythos.library.ucy.ac.cy/bitstream/handle/10797/14...

I am the first author.

Re: Classifying all of the pdfs on the internet

#64
post #58

Earlier quoted context omitted.

Also a neurodivergent person I feel very much discriminated against when a whole continent weaponizes the law to protect scam artists who weaponize their social skills to steal from people. It makes me feel unwelcome going to Europe and for all the handwriting about Europe’s poor economic performance it is yet another explanation of why Europe is falling behind — their wealth is being stolen by people who can’t be he…

Which scam artists are you referring to?

The ones who have filed lawsuits to try to get people in Europe to forget about their crimes.

Re: Classifying all of the pdfs on the internet

#65
post #59

I don’t have 8TB laying around, but we can be a bit more clever.... In particular I cared about a specific column called url. I really care about the urls because they essentially tell us a lot more from a website than what meats the eye. I'm I correct that it is only only using the URL of the PDF to do classification? Maybe still useful, but that's quite a different story than "classifying all the pdfs".

It’s just classifying the URLs if that’s the case.

The legwork to classify PDFs is already done, and the authorship of the article can go to anyone who can get a grant for a $400 NewEgg order for an 8TB drive.

Re: Classifying all of the pdfs on the internet

#66

Earlier quoted context omitted.

I haven't downvoted you but it is presumably because of your hasty typing or lack of proofreading/research. 33TB (first google result from 5 years ago) not 33GB. Larger figures from more recently.

>hasty typing or lack of proofreading/research This is exactly what I meant with "HN is becoming quite hostile" * I brought up something I looked up to support GP's argument. * The argument is correct. * I do it in good faith. * G is literally next to T. * I even praise the article, while at it. "Oh, but you made a typo!". Good luck, guys. I'm out. PS. I will give my whole 7 figure net worth, no questions asked, tran…

I haven't ever made a typo, all of my mispelings are intended and therefore not mistakes

Re: Classifying all of the pdfs on the internet

#67
post #62
post #5

Earlier quoted context omitted.

Yeah 8TB is really tiny. Google scholar was estimated to index 160.000.000 pdfs in 2015.[0] If we assume that a third of those are not behind paywalls, and average pdf size is 1mb, its ends up as something above 50TB of documents. Almost ten years later the number of available pdfs of just scholarly communication should be substantially higher. [0] https://link.springer.com/article/10.1007/s11192-015-1614-6

Anna's archive has some 300M pdfs.

We're talking about the open web here. But yeah that's the point, the dataset is unreasonably small.

Re: Classifying all of the pdfs on the internet

#68
This seems like cool work but with a ton of "marketing hype speak" that immediately gets watered down by the first paragraph.

Ordering of statements.

1. (Title) Classifying all of the pdfs on the internet

2. (First Paragraph) Well not all, but all the PDFs in Common Crawl

3. (First Image) Well not all of them, but 500k of them.

I am not knocking the project, but while categorizing 500k PDFs is something we couldnt necessarily do well a few years ago, this is far from "The internet's PDFs".

Re: Classifying all of the pdfs on the internet

#69

This seems like cool work but with a ton of "marketing hype speak" that immediately gets watered down by the first paragraph. Ordering of statements. 1. (Title) Classifying all of the pdfs on the internet 2. (First Paragraph) Well not all, but all the PDFs in Common Crawl 3. (First Image) Well not all of them, but 500k of them. I am not knocking the project, but while categorizing 500k PDFs is something we couldnt ne…

[deleted]

Re: Classifying all of the pdfs on the internet

#70

This seems like cool work but with a ton of "marketing hype speak" that immediately gets watered down by the first paragraph. Ordering of statements. 1. (Title) Classifying all of the pdfs on the internet 2. (First Paragraph) Well not all, but all the PDFs in Common Crawl 3. (First Image) Well not all of them, but 500k of them. I am not knocking the project, but while categorizing 500k PDFs is something we couldnt ne…

Overpromise with headline, underdeliver on details.
Post reply on HN