There are also other documents that appear to simulate a scanned document but completely lack the “real-world noise” expected with physical paper-based workflows. The much crisper images appear almost perfect without random artifacts or background noise, and with the exact same amount of image skew across multiple pages. Thanks to the borders around each page of text, page skew can easily be measured, such as with VOL00007\IMAGES\0001\EFTA00009229.pdf. It is highly likely these PDFs were created by rendering original content (from a digital document) to an image (e.g., via print to image or save to image functionality) and then applying image processing such as skew, downscaling, and color reduction.
A case study in PDF forensics: The Epstein PDFs
71–80 of 254 posts
Re: A case study in PDF forensics: The Epstein PDFs
#72Has anyone analysed JE's writing style and looked for matches in archived 4chan posts or content from similar platforms? Same with Ghislaine, there should be enough data to identify them atp right? I don't buy the MaxwellHill claims for various reasons but it doesn't mean there's nothing to find.
People always claimed this as a data leak vector but I've always been sceptical. Like just writing style and vocabulary is probably extremely shared among too many people to narrow it down much. (How people that you know could have written this reply?) The counter argument is that he had a very specific style in his mail so maybe this is a special case.
even when people deliberately try to feign some aspects (e.g. switching writing styles for different pseudonyms), they will almost always slip up and revert to their most comfortable style over time. which is great, because if they aren't also regularly changing pseudonyms (which are also subject to limited stylometry, so pseudonym creation should be somewhat randomized in name, location, etc.), you only need to catch them slipping once to get the whole history of that pseudonym (and potentially others, once that one is confirmed).
Re: A case study in PDF forensics: The Epstein PDFs
#73Has anyone analysed JE's writing style and looked for matches in archived 4chan posts or content from similar platforms? Same with Ghislaine, there should be enough data to identify them atp right? I don't buy the MaxwellHill claims for various reasons but it doesn't mean there's nothing to find.
Re: A case study in PDF forensics: The Epstein PDFs
#74Re: A case study in PDF forensics: The Epstein PDFs
#75Has anyone analysed JE's writing style and looked for matches in archived 4chan posts or content from similar platforms? Same with Ghislaine, there should be enough data to identify them atp right? I don't buy the MaxwellHill claims for various reasons but it doesn't mean there's nothing to find.
Stylometry is extremely sophisticated even with simple n-gram analysis. There's a demo of this that can easily pick out who you are on HN just based on a few paragraphs of your own writing, based on N-gram analysis. https://news.ycombinator.com/item?id=33755016 You can also unironically spot most types of AI writing this way. The approaches based on training another transformer to spot "AI generated" content are wron…
Re: A case study in PDF forensics: The Epstein PDFs
#76Earlier quoted context omitted.
I think some of the released documents included images of victims, which where redacted. So it's not necessarily malicious removals
That's my understanding too, so archiving the unredacted images could mean holding CSAM.
Re: A case study in PDF forensics: The Epstein PDFs
#77Earlier quoted context omitted.
Which meeting are you seeing? That search doesn't seem to work for me, I'm only seeing the one Jan 2012.
It doesn't show up in JMail for some reason, but it's this email: https://www.justice.gov/epstein/files/DataSet%2010/EFTA01852...
https://www.justice.gov/epstein/files/DataSet%2010/EFTA01992...
Re: A case study in PDF forensics: The Epstein PDFs
#78What is the legal basis for releasing the someone's private files and communications? If they can do it to Epstein, they can do it to you, to the Washington Post journalist, to former President Clinton, etc. Is the scope at least limited somehow? Generally I favor transparency, but of course probably the most important parts are withheld.
Re: A case study in PDF forensics: The Epstein PDFs
#79This is so incredibly useful to me right now for incidental reasons I am commenting to make sure I can get back to it.
Re: A case study in PDF forensics: The Epstein PDFs
#80I found this part interesting: There are also other documents that appear to simulate a scanned document but completely lack the “real-world noise” expected with physical paper-based workflows. The much crisper images appear almost perfect without random artifacts or background noise, and with the exact same amount of image skew across multiple pages. Thanks to the borders around each page of text, page skew can easi…
https://www.justice.gov/epstein/files/DataSet%207/EFTA000092...