>
someone made a great ... in-memory low level pdf reading and writing data structureAre you suggesting Adobe's Core Object Application Programming Interface (COAPI) for PDF isn't sufficient?
Kidding!
I worked on print production software in the '90s. Stuff like image positioning (eg bookwork), trapping, color separations, etc. Adobe's SDKs, for both PostScript and PDF, were most turrible. For our greenfield product for packaging (printing boxes), I wrote a minimalist PDF library, supporting just the feature set we needed. So simple.
Of course, PDF is now an ever growing katamari style All The Things amalgamation of, oops, sorry I ran out of adjectives.
Back to your point: after URLs and HTTP, the DOM is the 3rd best thing spawned by "the web".
The DOM concept itself. Isomorphism between in-memory and serialized. That its all just an object graph. Composition over inheritance.
Not the actual DOM API; gods no.
I understand that API design is wicked hard. But how is it that of the Java tools, only JDOM2 (the sequel) managed to get the class hierarchy correct? So that incorrect usage is not permitted?
(I haven't looked at popular libraries for other languages. I assume they all also fell into the trap of transliterating JavaScript's DOM's API. Like dom4j and successors did.)
I'm just repeating your point (I think) that Adobe should have staked a strong starting conceptual position on PDF internals, what a PDF is. Something more WinForms and less Win32.
30+ (?!) years later, I'm still flubbergasted by PDF's success, despite Adobe's stewardship.
PS- And another thing...
For a print description language, I greatly preferred HP's PCL-5. Emotionally, it just feels more honest somehow. Initially, Adobe couldn't decide if PDF was for print control or documents. Customers wanted documents, so Adobe grudgingly complied, haphazardly.
At least "the web" had/has committees.