Interesting read, I did not know about Common Crawl. I feel like RTBF is kind of a lost battle these days with more and more crawlers for AI and whatnot. Once on the internet there is no way back, for better or for worse. This tangent aside, 8TB is really not a lot of data, it's just 8 consumer-grade 1TB hard drives. I find it hard to believe this is "the largest corpus of PDFs online", maybe the largest public one.…
Classifying all of the pdfs on the internet
71–80 of 117 posts
Re: Classifying all of the pdfs on the internet
#72I don’t have 8TB laying around, but we can be a bit more clever.... In particular I cared about a specific column called url. I really care about the urls because they essentially tell us a lot more from a website than what meats the eye. I'm I correct that it is only only using the URL of the PDF to do classification? Maybe still useful, but that's quite a different story than "classifying all the pdfs".
Re: Classifying all of the pdfs on the internet
#73Interesting read, I did not know about Common Crawl. I feel like RTBF is kind of a lost battle these days with more and more crawlers for AI and whatnot. Once on the internet there is no way back, for better or for worse. This tangent aside, 8TB is really not a lot of data, it's just 8 consumer-grade 1TB hard drives. I find it hard to believe this is "the largest corpus of PDFs online", maybe the largest public one.…
Re: Classifying all of the pdfs on the internet
#74Earlier quoted context omitted.
>hasty typing or lack of proofreading/research This is exactly what I meant with "HN is becoming quite hostile" * I brought up something I looked up to support GP's argument. * The argument is correct. * I do it in good faith. * G is literally next to T. * I even praise the article, while at it. "Oh, but you made a typo!". Good luck, guys. I'm out. PS. I will give my whole 7 figure net worth, no questions asked, tran…
Like I said, I didn't downvote and took the time to answer your question. I didn't take the time to sugarcoat it. You are interpreting bluntness as hostility; that's ultimately an issue for you to resolve.
You just have to read this site's guidelines and follow them.
Ez pz.
Re: Classifying all of the pdfs on the internet
#75Re: Classifying all of the pdfs on the internet
#76Re: Classifying all of the pdfs on the internet
#77Earlier quoted context omitted.
Like I said, I didn't downvote and took the time to answer your question. I didn't take the time to sugarcoat it. You are interpreting bluntness as hostility; that's ultimately an issue for you to resolve.
You don't have to sugarcoat it. You just have to read this site's guidelines and follow them. Ez pz.
Re: Classifying all of the pdfs on the internet
#78Classification is just a start. Wondering if it's worth doing something more -- like turning all of the text into Markdown or HTML? Would anyone find that interesting?
Re: Classifying all of the pdfs on the internet
#79Earlier quoted context omitted.
Which scam artists are you referring to?
The ones who have filed lawsuits to try to get people in Europe to forget about their crimes.
Re: Classifying all of the pdfs on the internet
#80Ive been playing with https://www.aryn.ai/ for Partitioning. Curious if anyone has tried these tools for better data extraction from PDFs. Any other suggestions? (I'm a bit disappointed that most of the discussion is about estimating the size of PDFs on the internet, I'd love to hear more about different approaches to extracting better data from the PDFs.)
Full disclosure: I'm an employee