Classifying all of the pdfs on the internet
51–60 of 117 posts
Re: Classifying all of the pdfs on the internet
#52Back in 2006 there were multiple 1tb collections of textbooks as torrents. I imagine the size and number has only grown since then.
The main difference were sites like chegg and many other sites started slurping them up to resell in some way.
Re: Classifying all of the pdfs on the internet
#53Earlier quoted context omitted.
I haven't downvoted you but it is presumably because of your hasty typing or lack of proofreading/research. 33TB (first google result from 5 years ago) not 33GB. Larger figures from more recently.
>hasty typing or lack of proofreading/research This is exactly what I meant with "HN is becoming quite hostile" * I brought up something I looked up to support GP's argument. * The argument is correct. * I do it in good faith. * G is literally next to T. * I even praise the article, while at it. "Oh, but you made a typo!". Good luck, guys. I'm out. PS. I will give my whole 7 figure net worth, no questions asked, tran…
You are interpreting bluntness as hostility; that's ultimately an issue for you to resolve.
Re: Classifying all of the pdfs on the internet
#54Re: Classifying all of the pdfs on the internet
#55Re: Classifying all of the pdfs on the internet
#56Interesting read, I did not know about Common Crawl. I feel like RTBF is kind of a lost battle these days with more and more crawlers for AI and whatnot. Once on the internet there is no way back, for better or for worse. This tangent aside, 8TB is really not a lot of data, it's just 8 consumer-grade 1TB hard drives. I find it hard to believe this is "the largest corpus of PDFs online", maybe the largest public one.…
RTBF was a ludicrous concept before AI and these new crawlers. Only EU bureaucracts would have the hubris to believe you could actually, comprehensively remove information from the Internet. Once something is spread, it is there, forever.
Re: Classifying all of the pdfs on the internet
#57Earlier quoted context omitted.
RTBF was a ludicrous concept before AI and these new crawlers. Only EU bureaucracts would have the hubris to believe you could actually, comprehensively remove information from the Internet. Once something is spread, it is there, forever.
RTBF isn't about having your information wiped from the internet. Its a safe assumption any public information about you is completely out of your control as soon as its public. RTBF is about getting companies to get rid of any trace of you so they cannot use that data, not removing all traces about you across the internet.
your take is misleading enough to be considered wrong. It's "don't use public information about me in search engines, I don't want people to find that information about me", not simply "don't use my information for marketing purposes"
https://en.wikipedia.org/wiki/Right_to_be_forgotten
first paragraph of the article: The right to be forgotten (RTBF) is the right to have private information about a person be removed from Internet searches and other directories in some circumstances. The issue has arisen from desires of individuals to "determine the development of their life in an autonomous way, without being perpetually or periodically stigmatized as a consequence of a specific action performed in the past". The right entitles a person to have data about them deleted so that it can no longer be discovered by third parties, particularly through search engines.
Re: Classifying all of the pdfs on the internet
#58Earlier quoted context omitted.
RTBF was a ludicrous concept before AI and these new crawlers. Only EU bureaucracts would have the hubris to believe you could actually, comprehensively remove information from the Internet. Once something is spread, it is there, forever.
Also a neurodivergent person I feel very much discriminated against when a whole continent weaponizes the law to protect scam artists who weaponize their social skills to steal from people. It makes me feel unwelcome going to Europe and for all the handwriting about Europe’s poor economic performance it is yet another explanation of why Europe is falling behind — their wealth is being stolen by people who can’t be he…
Re: Classifying all of the pdfs on the internet
#59 I don’t have 8TB laying around, but we can be a bit more clever.... In particular I cared about a specific column called url. I really care about the urls because they essentially tell us a lot more from a website than what meats the eye.
I'm I correct that it is only only using the URL of the PDF to do classification? Maybe still useful, but that's quite a different story than "classifying all the pdfs".Re: Classifying all of the pdfs on the internet
#60Earlier quoted context omitted.
I haven't downvoted you but it is presumably because of your hasty typing or lack of proofreading/research. 33TB (first google result from 5 years ago) not 33GB. Larger figures from more recently.
>hasty typing or lack of proofreading/research This is exactly what I meant with "HN is becoming quite hostile" * I brought up something I looked up to support GP's argument. * The argument is correct. * I do it in good faith. * G is literally next to T. * I even praise the article, while at it. "Oh, but you made a typo!". Good luck, guys. I'm out. PS. I will give my whole 7 figure net worth, no questions asked, tran…