I wonder if we're doing the wrong thing blocking them with invasive tools like cloudflare? If all you're concerned about is server load, wouldn't it be better to just offer a tar file containing all of your pages they can download instead? The models are months out of date, so a monthly dumb would surely satisfy them. There could even be some coordination for this. They're going to crawl anyway. We can either coopera…
If the crawlers were aware of these archive files, and would be willing to use it, then that would help, but it isn't. (It would also help to know which dynamic files are worthless for archiving and mirroring, but they will often ignore that.)