Live data from Hacker News

A case study in PDF forensics: The Epstein PDFs

pdfa.org

91–100 of 254 posts

Re: A case study in PDF forensics: The Epstein PDFs

#91

Has anyone analysed JE's writing style and looked for matches in archived 4chan posts or content from similar platforms? Same with Ghislaine, there should be enough data to identify them atp right? I don't buy the MaxwellHill claims for various reasons but it doesn't mean there's nothing to find.

Stylometry is extremely sophisticated even with simple n-gram analysis. There's a demo of this that can easily pick out who you are on HN just based on a few paragraphs of your own writing, based on N-gram analysis. https://news.ycombinator.com/item?id=33755016 You can also unironically spot most types of AI writing this way. The approaches based on training another transformer to spot "AI generated" content are wron…

Hacker News is one of the best places for this, because people write relatively long posts and generally try to have novel ideas. On 4chan, most posts are very short memey quips, so everybody's style is closer to each others than it is to their normal writing style.

Re: A case study in PDF forensics: The Epstein PDFs

#92
post #39

Earlier quoted context omitted.

He met with moot ("he is sensitive, be gentile", search on jmail), and within a few days the /pol/ board got created, starting a culture war in the US, leading to Trump getting elected president. Absolutely nuts.

Few thoughts: in context it's not nuts at all: - moot was fundraising for his VC backed startup during the years the emails are in, and he was likely connected via mutuals in USV or other firms. These meetings were clearly around him trying to solicit investment in his canv.as project. - /pol/ was /new/ being returned; the ethos of the board had already existed for a long time and the decision to undo the deletion of…

Besides /new/ there was also /n/ (not at that time about transportation.) Moot's war with people being racist on 4chan had many back and forths before /pol/ was created.

Re: A case study in PDF forensics: The Epstein PDFs

#93
post #14

> DoJ explicitly avoids JPEG images in the PDFs probably because they appreciate that JPEGs often contain identifiable information, such as EXIF, IPTC, or XMP metadata Maybe I'm underestimating the issue at full, but isn't this a very lightweight problem to solve? Is converting the images to lower DPI formats/versions really any easier than just stripping the metadata? Surely the DOJ and similar justice agencies have…

Image metadata is the wild west of structured text. The developer of the foremost tool for dealing with it (exiftool) has made 'remove metadata' feature but still disclaims that it is not able to remove everything.

Re: A case study in PDF forensics: The Epstein PDFs

#94
post #62
post #39

Earlier quoted context omitted.

He met with moot ("he is sensitive, be gentile", search on jmail), and within a few days the /pol/ board got created, starting a culture war in the US, leading to Trump getting elected president. Absolutely nuts.

Just to substantiate this a bit: I remember a gleeful consensus in certain circles being that /pol/ and /r/the_donald had "memed Trump into the White House". It's much more complicated than that, but there's certainly an element of truth there.

Then Reddit and almost all of social media went on to purge trump and pro trump content. The Donald was banned. Trump deplatformed across social media.

Re: A case study in PDF forensics: The Epstein PDFs

#95
post #14

> DoJ explicitly avoids JPEG images in the PDFs probably because they appreciate that JPEGs often contain identifiable information, such as EXIF, IPTC, or XMP metadata Maybe I'm underestimating the issue at full, but isn't this a very lightweight problem to solve? Is converting the images to lower DPI formats/versions really any easier than just stripping the metadata? Surely the DOJ and similar justice agencies have…

Image metadata is the wild west of structured text. The developer of the foremost tool for dealing with it (exiftool) has made 'remove metadata' feature but still disclaims that it is not able to remove everything.

How could that be possible? Isn't JPEG a fairly straightforward container for JFIF+metadata?

Re: A case study in PDF forensics: The Epstein PDFs

#96
post #44

What is the legal basis for releasing the someone's private files and communications? If they can do it to Epstein, they can do it to you, to the Washington Post journalist, to former President Clinton, etc. Is the scope at least limited somehow? Generally I favor transparency, but of course probably the most important parts are withheld.

https://www.congress.gov/bill/119th-congress/house-bill/4405...

I assume this could not have passed while he was alive, because of the "bill of attainder" thing?

(It also surprises me that this passed anyway, given that both sides of the aisle seem to have people with clear reason to keep it covered up... ?)

(Also, Maxwell is specifically named, and is still alive... ?)

Re: A case study in PDF forensics: The Epstein PDFs

#97
post #87

Earlier quoted context omitted.

Not that I can speak from personal experience or anything... But somebody on an email chain may have requested a scanned version of the document to ensure there is no metadata and the employee might have found it easier to just flatten the pdf and apply a graphical filter to make the document appear like a scanned document. There might even be a webtool available somewhere to do so, I wouldn't know...

> the employee might have found it easier to just flatten the pdf and apply a graphical filter to make the document appear like a scanned document Is that remotely plausible? I can't imaging faking a scan being easier than just walking down the hall to the copier room.

Depending on their technical capability, yes.

I mean even in this thread you got what are essentially one-liners to do it.

Definitely less hassle then doing it irl

Re: A case study in PDF forensics: The Epstein PDFs

#99
post #87

Earlier quoted context omitted.

Very interesting. That document in particular seems to be an interview of A. Acosta by the DoJ from 2019. But what reason would the FBI have for pretending it's a scanned document, if it is genuine? Perhaps there's some aspect of Epstein's deal with Acosta that they'd rather not reveal to the public? https://www.justice.gov/epstein/files/DataSet%207/EFTA000092...

Not that I can speak from personal experience or anything... But somebody on an email chain may have requested a scanned version of the document to ensure there is no metadata and the employee might have found it easier to just flatten the pdf and apply a graphical filter to make the document appear like a scanned document. There might even be a webtool available somewhere to do so, I wouldn't know...

[dead]

Re: A case study in PDF forensics: The Epstein PDFs

#100
post #58

Earlier quoted context omitted.

People always claimed this as a data leak vector but I've always been sceptical. Like just writing style and vocabulary is probably extremely shared among too many people to narrow it down much. (How people that you know could have written this reply?) The counter argument is that he had a very specific style in his mail so maybe this is a special case.

If you have a large enough set to test against and a specific person you are looking for, this is totally doable currently.

Of course it's doable. The question is how reliable the results are.
Post reply on HN