Live data from Hacker News

A case study in PDF forensics: The Epstein PDFs

pdfa.org

121–130 of 254 posts

Re: A case study in PDF forensics: The Epstein PDFs

#121
post #97

Earlier quoted context omitted.

> the employee might have found it easier to just flatten the pdf and apply a graphical filter to make the document appear like a scanned document Is that remotely plausible? I can't imaging faking a scan being easier than just walking down the hall to the copier room.

Depending on their technical capability, yes. I mean even in this thread you got what are essentially one-liners to do it. Definitely less hassle then doing it irl

I know I'm not the brightest bulb by any measure, but do some people really take less than at least a few minutes to come up with one-liners for problems as novel as graphical transformations to PDFs? Maybe if the presumed techie hacker / federal worker took it as an amusing challenge I could see this being done, but genuinely out of pure laziness? That's incredible if true.

Re: A case study in PDF forensics: The Epstein PDFs

#123
post #84

Earlier quoted context omitted.

I think the logic is Pol didn't need to reach the masses, the masses only consume content they don't create it. You only need to radicalize the few people who then go on to be the 1% of people commenting and posting.

There's an old joke that 9gag* only reposts stuff from Reddit and Reddit only reposts stuff from 4chan and 4chan is the origin of all meme culture. This joke was widespread enough to reach myself and my friend group back in the day, even though none of used 4chan or Reddit. If you radicalise the 0.01% of people who are prolific meme creators, you radicalise the masses. * I did say old...

And Facebook repeats stuff from 9gag

Re: A case study in PDF forensics: The Epstein PDFs

#125

Earlier quoted context omitted.

Haven't seen anything particular about that, but lots of the documents with names that were half-redacted contain OCRd text that is completely garbled, but olmocr-2-7b seems to handle it just fine. Unsure if they just had sucky processes or if there is something else going on.

Might be a good fit for uploading a git repo and crowdsourcing

GitHub would ban you

Re: A case study in PDF forensics: The Epstein PDFs

#126
post #87

Earlier quoted context omitted.

Not that I can speak from personal experience or anything... But somebody on an email chain may have requested a scanned version of the document to ensure there is no metadata and the employee might have found it easier to just flatten the pdf and apply a graphical filter to make the document appear like a scanned document. There might even be a webtool available somewhere to do so, I wouldn't know...

> the employee might have found it easier to just flatten the pdf and apply a graphical filter to make the document appear like a scanned document Is that remotely plausible? I can't imaging faking a scan being easier than just walking down the hall to the copier room.

It's thousands of pages, surely investing some time in a script is faster. They were in a rush as well.

If they were faking the documents rather than the delivery method they definitely could have invested some time in flawless looks.

Re: A case study in PDF forensics: The Epstein PDFs

#127
post #87

Earlier quoted context omitted.

Not that I can speak from personal experience or anything... But somebody on an email chain may have requested a scanned version of the document to ensure there is no metadata and the employee might have found it easier to just flatten the pdf and apply a graphical filter to make the document appear like a scanned document. There might even be a webtool available somewhere to do so, I wouldn't know...

> the employee might have found it easier to just flatten the pdf and apply a graphical filter to make the document appear like a scanned document Is that remotely plausible? I can't imaging faking a scan being easier than just walking down the hall to the copier room.

The time advantage of faking a scan becomes better the more pages you have to scan.

https://xkcd.com/1205/

Re: A case study in PDF forensics: The Epstein PDFs

#128
post #95

Earlier quoted context omitted.

Image metadata is the wild west of structured text. The developer of the foremost tool for dealing with it (exiftool) has made 'remove metadata' feature but still disclaims that it is not able to remove everything.

How could that be possible? Isn't JPEG a fairly straightforward container for JFIF+metadata?

"Fairly straightforward" is incorrect. Not an authority to describe in more detail, but the most tricky blocker I'm aware of are these proprietary "MakerNote" tags from camera manufacturers, which are (often undocumented) binary blobs. exiftool might not even know what's in there, let alone how to safely remove it without corrupting the file.

Re: A case study in PDF forensics: The Epstein PDFs

#129
post #39

Earlier quoted context omitted.

I'm pretty sure Epstein tried to meet with moot at least once: https://www.jmail.world/search?q=chris+poole

He met with moot ("he is sensitive, be gentile", search on jmail), and within a few days the /pol/ board got created, starting a culture war in the US, leading to Trump getting elected president. Absolutely nuts.

be gentile

We're just not going to talk about that one I suppose?

Re: A case study in PDF forensics: The Epstein PDFs

#130

Earlier quoted context omitted.

> the employee might have found it easier to just flatten the pdf and apply a graphical filter to make the document appear like a scanned document Is that remotely plausible? I can't imaging faking a scan being easier than just walking down the hall to the copier room.

It's thousands of pages, surely investing some time in a script is faster. They were in a rush as well. If they were faking the documents rather than the delivery method they definitely could have invested some time in flawless looks.

Or more-realistic flawed looks as the case is here.
Post reply on HN