Live data from Hacker News

A case study in PDF forensics: The Epstein PDFs

pdfa.org

181–190 of 254 posts

Re: A case study in PDF forensics: The Epstein PDFs

#181

I found this part interesting: There are also other documents that appear to simulate a scanned document but completely lack the “real-world noise” expected with physical paper-based workflows. The much crisper images appear almost perfect without random artifacts or background noise, and with the exact same amount of image skew across multiple pages. Thanks to the borders around each page of text, page skew can easi…

The real question is: Which of the documents are the ones that are "simulating" scanned documents, and what political narrative do they reinforce? The only reason I can think of for why someone would want to do this is to pass off fraudulent or AI generated images as real.

This. Slip in a few thousand “fakes” with the trove of goods to be able to fabricate a narrative.

Re: A case study in PDF forensics: The Epstein PDFs

#182

I found this part interesting: There are also other documents that appear to simulate a scanned document but completely lack the “real-world noise” expected with physical paper-based workflows. The much crisper images appear almost perfect without random artifacts or background noise, and with the exact same amount of image skew across multiple pages. Thanks to the borders around each page of text, page skew can easi…

Such a weird way to do it when it would be a vastly easier to just blow the document out to paper and re-scan it.

Re: A case study in PDF forensics: The Epstein PDFs

#183
post #182

I found this part interesting: There are also other documents that appear to simulate a scanned document but completely lack the “real-world noise” expected with physical paper-based workflows. The much crisper images appear almost perfect without random artifacts or background noise, and with the exact same amount of image skew across multiple pages. Thanks to the borders around each page of text, page skew can easi…

Such a weird way to do it when it would be a vastly easier to just blow the document out to paper and re-scan it.

Vastly easier when you do it to one or a handful of documents.

But if you want to do it to 2000 documents...

Re: A case study in PDF forensics: The Epstein PDFs

#184

Earlier quoted context omitted.

>It makes me wanna write an AI browser assistant that can take my comments and stylize them randomly to make it harder to use these sorts of forensics against me The old trick years ago was to translate from English to different language and back (possibly repeating). I'd be curious how helpful it is against stylometry detection?

The old trick years ago was to translate from English to different language and back (possibly repeating). I'd be curious how helpful it is against stylometry detection? If you want to be grouped with foreigners who don't know English, it might work well, although word choices may still be distinctive enough to differentiate even when translated.

Assuming the source language is English, going to a romance language and back wouldn't be too hard grammar wise, but could easily wipe out a lot of non-Latin-descended words if you use the right approach to translation.

Re: A case study in PDF forensics: The Epstein PDFs

#186

Earlier quoted context omitted.

Pretty devious tactic if so. Chilling effect on both any further witnesses and anybody interested in archiving the data (gives them an ethical conundrum at least). In addition to giving them (the feds) a convenient excuse to take down random docs.

Not about a ethical conondrum when rehosting. Anyone who rehosts the whole files can be accused of hosting child porn and doxxing and taken down.

Very convenient for Epstein and his associates

Re: A case study in PDF forensics: The Epstein PDFs

#187

Earlier quoted context omitted.

You think the personal lawyers of Donald Trump Pam Bondi and Todd Blanche will follow the money unbiasedly? As well as children's book, The Plot Against the King, author Kash Patel and FBI director? As well as Russian asset herself, Tulsa Gabbard director of National Intelligence want to do anything against their power source?

The democrats had these files and all that information in their power for what? 5 years? And what did they do? Stop making this a partisan issue. It’s not, and nobody that’s not completely biased beyond any rationality will ever see it as such.

Democrats do nothing because they are useless. Republicans do nothing because they are implicated.

Re: A case study in PDF forensics: The Epstein PDFs

#188
post #145

Earlier quoted context omitted.

Don't throw the baby out with the bathwater. "Ooh, la" sounds really unnatural. But on a serious note, what did "la" mean in your context? I've never seen this.

It’s a common thing for speakers of Singaporean English to end sentences with la/leh. But no idea if that’s what’s going on here.

In one use case, it is kind of a verbal exclamation point, but it has more meanings and uses than just that. Likely originates from Hokkien, but it has evolved into it is own thing. If you are curious, more details here https://en.wikipedia.org/wiki/Singlish

Re: A case study in PDF forensics: The Epstein PDFs

#189
Stylometry works. I've seen it used it cases where the individual was identified from a group.

One thing that is telling about the Epstein case study is how long it has stayed in public view. Pizzagate, which involved more powerful people, was shut down faster than I've ever seen for anything else. I still remember and have archived the more extreme content it's sick.

Re: A case study in PDF forensics: The Epstein PDFs

#190

Earlier quoted context omitted.

> the employee might have found it easier to just flatten the pdf and apply a graphical filter to make the document appear like a scanned document Is that remotely plausible? I can't imaging faking a scan being easier than just walking down the hall to the copier room.

Working from home and no scanner in the house?

No printer.
Post reply on HN