Interesting read, I did not know about Common Crawl. I feel like RTBF is kind of a lost battle these days with more and more crawlers for AI and whatnot. Once on the internet there is no way back, for better or for worse. This tangent aside, 8TB is really not a lot of data, it's just 8 consumer-grade 1TB hard drives. I find it hard to believe this is "the largest corpus of PDFs online", maybe the largest public one.…
> RTBF Right to be forgotten, not the Belgian public service broadcaster ( https://en.wikipedia.org/wiki/RTBF )?
Classifying all of the pdfs on the internet
61–70 of 117 posts
Re: Classifying all of the pdfs on the internet
#62Interesting read, I did not know about Common Crawl. I feel like RTBF is kind of a lost battle these days with more and more crawlers for AI and whatnot. Once on the internet there is no way back, for better or for worse. This tangent aside, 8TB is really not a lot of data, it's just 8 consumer-grade 1TB hard drives. I find it hard to believe this is "the largest corpus of PDFs online", maybe the largest public one.…
Yeah 8TB is really tiny. Google scholar was estimated to index 160.000.000 pdfs in 2015.[0] If we assume that a third of those are not behind paywalls, and average pdf size is 1mb, its ends up as something above 50TB of documents. Almost ten years later the number of available pdfs of just scholarly communication should be substantially higher. [0] https://link.springer.com/article/10.1007/s11192-015-1614-6
Re: Classifying all of the pdfs on the internet
#63Newton, G., A. Callahan & M. Dumontier. 2009. Semantic Journal Mapping for Search Visualization in a Large Scale Article Digital Library. Second Workshop on Very Large Digital Libraries at the European Conference on Digital Libraries (ECDL) 2009. https://lekythos.library.ucy.ac.cy/bitstream/handle/10797/14...
I am the first author.
Re: Classifying all of the pdfs on the internet
#64Earlier quoted context omitted.
Also a neurodivergent person I feel very much discriminated against when a whole continent weaponizes the law to protect scam artists who weaponize their social skills to steal from people. It makes me feel unwelcome going to Europe and for all the handwriting about Europe’s poor economic performance it is yet another explanation of why Europe is falling behind — their wealth is being stolen by people who can’t be he…
Which scam artists are you referring to?
Re: Classifying all of the pdfs on the internet
#65I don’t have 8TB laying around, but we can be a bit more clever.... In particular I cared about a specific column called url. I really care about the urls because they essentially tell us a lot more from a website than what meats the eye. I'm I correct that it is only only using the URL of the PDF to do classification? Maybe still useful, but that's quite a different story than "classifying all the pdfs".
The legwork to classify PDFs is already done, and the authorship of the article can go to anyone who can get a grant for a $400 NewEgg order for an 8TB drive.
Re: Classifying all of the pdfs on the internet
#66Earlier quoted context omitted.
I haven't downvoted you but it is presumably because of your hasty typing or lack of proofreading/research. 33TB (first google result from 5 years ago) not 33GB. Larger figures from more recently.
>hasty typing or lack of proofreading/research This is exactly what I meant with "HN is becoming quite hostile" * I brought up something I looked up to support GP's argument. * The argument is correct. * I do it in good faith. * G is literally next to T. * I even praise the article, while at it. "Oh, but you made a typo!". Good luck, guys. I'm out. PS. I will give my whole 7 figure net worth, no questions asked, tran…
Re: Classifying all of the pdfs on the internet
#67Earlier quoted context omitted.
Yeah 8TB is really tiny. Google scholar was estimated to index 160.000.000 pdfs in 2015.[0] If we assume that a third of those are not behind paywalls, and average pdf size is 1mb, its ends up as something above 50TB of documents. Almost ten years later the number of available pdfs of just scholarly communication should be substantially higher. [0] https://link.springer.com/article/10.1007/s11192-015-1614-6
Anna's archive has some 300M pdfs.
Re: Classifying all of the pdfs on the internet
#68Ordering of statements.
1. (Title) Classifying all of the pdfs on the internet
2. (First Paragraph) Well not all, but all the PDFs in Common Crawl
3. (First Image) Well not all of them, but 500k of them.
I am not knocking the project, but while categorizing 500k PDFs is something we couldnt necessarily do well a few years ago, this is far from "The internet's PDFs".
Re: Classifying all of the pdfs on the internet
#69This seems like cool work but with a ton of "marketing hype speak" that immediately gets watered down by the first paragraph. Ordering of statements. 1. (Title) Classifying all of the pdfs on the internet 2. (First Paragraph) Well not all, but all the PDFs in Common Crawl 3. (First Image) Well not all of them, but 500k of them. I am not knocking the project, but while categorizing 500k PDFs is something we couldnt ne…
Re: Classifying all of the pdfs on the internet
#70This seems like cool work but with a ton of "marketing hype speak" that immediately gets watered down by the first paragraph. Ordering of statements. 1. (Title) Classifying all of the pdfs on the internet 2. (First Paragraph) Well not all, but all the PDFs in Common Crawl 3. (First Image) Well not all of them, but 500k of them. I am not knocking the project, but while categorizing 500k PDFs is something we couldnt ne…