Earlier quoted context omitted.
Cool. Let me know when I can extract the full text contents from such files using common, built-in tools on your average fresh install of MacOS/Linux. PDF works just fine. It presents a feasible view of the original data, and allows for data harvesting with ease.
Let me know when I can extract the full text You can extract the full text from these (with whatever tools you like) with better fidelity than you can from a pdf, which is a lossy conversion from the same source. This seems to barely merit debating, unless I'm missing something. PDF works just fine. I'm sure it works for you and I'm not harbouring any delusions I'm going to talk you out of your decades-established wo…
I am not finding this to be true. Pretty much every PDF I have has been usable for extracting the text content - unless the Web authors intentionally work to obfuscate/disable this functionality, i.e. using images to display text content.
>PDF is not a good way to archive web pages, either manually or programmatically.
I disagree, entirely, with your conclusion - you haven't made a strong argument. 20,000+ fully-searchable, indexable, accessible-in-offline PDF files vs. your opinion so far. I don't see any of the issues you've stated are insurmountable - in fact, I find the reality to be completely the opposite to your stated opinion. Please expand on this if you have the energy.