Earlier quoted context omitted.
I get the use case, but in most cases (and particularly this one) I'm sure it would be much better to implement that client-side. You may have seen in the WARC standard that they already do de-duplication based on hashes and use pointers after the first store. So this is exactly a case where FS-level dedup is not all that good.
WARC only does deduping within a single WARC, I'm talking about deduping across millions of WARCs.
[edit] Should add a link, this is a pretty good overview, but you can also look at implementations such as the new zeno crawler.
https://support.archive-it.org/hc/en-us/articles/208001016-A...