Live data from Hacker News

Classifying all of the pdfs on the internet

snats.xyz

71–80 of 117 posts

Re: Classifying all of the pdfs on the internet

#71
post #2

Interesting read, I did not know about Common Crawl. I feel like RTBF is kind of a lost battle these days with more and more crawlers for AI and whatnot. Once on the internet there is no way back, for better or for worse. This tangent aside, 8TB is really not a lot of data, it's just 8 consumer-grade 1TB hard drives. I find it hard to believe this is "the largest corpus of PDFs online", maybe the largest public one.…

[deleted]

Re: Classifying all of the pdfs on the internet

#72
post #59

I don’t have 8TB laying around, but we can be a bit more clever.... In particular I cared about a specific column called url. I really care about the urls because they essentially tell us a lot more from a website than what meats the eye. I'm I correct that it is only only using the URL of the PDF to do classification? Maybe still useful, but that's quite a different story than "classifying all the pdfs".

[deleted]

Re: Classifying all of the pdfs on the internet

#73
post #2

Interesting read, I did not know about Common Crawl. I feel like RTBF is kind of a lost battle these days with more and more crawlers for AI and whatnot. Once on the internet there is no way back, for better or for worse. This tangent aside, 8TB is really not a lot of data, it's just 8 consumer-grade 1TB hard drives. I find it hard to believe this is "the largest corpus of PDFs online", maybe the largest public one.…

[deleted]

Re: Classifying all of the pdfs on the internet

#74

Earlier quoted context omitted.

>hasty typing or lack of proofreading/research This is exactly what I meant with "HN is becoming quite hostile" * I brought up something I looked up to support GP's argument. * The argument is correct. * I do it in good faith. * G is literally next to T. * I even praise the article, while at it. "Oh, but you made a typo!". Good luck, guys. I'm out. PS. I will give my whole 7 figure net worth, no questions asked, tran…

Like I said, I didn't downvote and took the time to answer your question. I didn't take the time to sugarcoat it. You are interpreting bluntness as hostility; that's ultimately an issue for you to resolve.

You don't have to sugarcoat it.

You just have to read this site's guidelines and follow them.

Ez pz.

Re: Classifying all of the pdfs on the internet

#77

Earlier quoted context omitted.

Like I said, I didn't downvote and took the time to answer your question. I didn't take the time to sugarcoat it. You are interpreting bluntness as hostility; that's ultimately an issue for you to resolve.

You don't have to sugarcoat it. You just have to read this site's guidelines and follow them. Ez pz.

Have been throughout. Anyway, I hope you are able to reconsider and move on within HN.

Re: Classifying all of the pdfs on the internet

#78

Classification is just a start. Wondering if it's worth doing something more -- like turning all of the text into Markdown or HTML? Would anyone find that interesting?

There are a lot of webcrawlers where the chief feature is turning the website into markdown, I don't quite understand what they are doing for me thats useful since I can just do something like `markdownify(my_html)` or whatever, all this to say is that I wouldn't find this useful, but also clearly people think this is a useful feature as part of an LLM pipeline.

Re: Classifying all of the pdfs on the internet

#79
post #58

Earlier quoted context omitted.

Which scam artists are you referring to?

The ones who have filed lawsuits to try to get people in Europe to forget about their crimes.

Do you have some examples? I was not aware that this was a thing. And are we talking about sentences fully served, or before that time?

Re: Classifying all of the pdfs on the internet

#80

Ive been playing with https://www.aryn.ai/ for Partitioning. Curious if anyone has tried these tools for better data extraction from PDFs. Any other suggestions? (I'm a bit disappointed that most of the discussion is about estimating the size of PDFs on the internet, I'd love to hear more about different approaches to extracting better data from the PDFs.)

https://www.sensible.so/

Full disclosure: I'm an employee

Post reply on HN