Be careful with PDF! There are many ambiguities in the specification that are implemented differently between parsers, as well as implicitly accepted malformations that almost all parsers will silently accept without warning. It is very easy to accidentally produce so-called file format schizophrenia: When the same file is rendered differently between two parsers. For example, with PDF, what if you have a PDF object…
In the README for that repository it mentions "schizophrenic files". What is a schizophrenic file, out of interest?
Here's a CCC talk on it: https://media.ccc.de/v/MRMCD2014_-_6008_-_en_-_grossbaustell...
And the slies from the talk: https://www.slideshare.net/ange4771/schizophrenic-files-v2