Sorry if i don't understand this totally, but why is scraping for pdf and ppt/pptx files and mirroring them illegal? If you can reach that files just scraping it means they are somehow open to public access. No joke, i am genuinely asking.
Scraping, mostly no. Using them to make profit: yes. Since you then violate the copyright of the author as the original action was solely to make it public (assuming no profit was intended). edit: Rehosting is not allowed as far as I know in Dutch law if you are attempting to make profit of it by not requesting it from the original owner. Rehosting it and not taking advantage and linking to the original article is al…
Making a file public without restriction doesn’t mean another can’t make profit (my ISP makes a profit by transmitting the file to me; gmail makes profit when I email the file to me; google makes a profit when they cache the file; archive makes a profit when they archive a file; etc etc).
I think if this were non-public docs then the case is clearer. But by releasing a document publicly with unlimited access via URL the author explicitly allows unlimited distribution (and due to the nature of tcp/ip redistribution). If an author wants to restrict distribution then they should restrict distribution using available protocols.