Live data from Hacker News

A case study in PDF forensics: The Epstein PDFs

pdfa.org

131–140 of 254 posts

Re: A case study in PDF forensics: The Epstein PDFs

#131

Earlier quoted context omitted.

If you have a large enough set to test against and a specific person you are looking for, this is totally doable currently.

Of course it's doable. The question is how reliable the results are.

It just needs to find the needles in the haystack. Humans can better verify if they're truly needles.

Re: A case study in PDF forensics: The Epstein PDFs

#132
post #103

Earlier quoted context omitted.

People do change over time, I used to write "ha" after every sentence for some reason

You know, i had a particularly cringy period in which i put "la" at the end of sentences.

Don't throw the baby out with the bathwater. "Ooh, la" sounds really unnatural.

But on a serious note, what did "la" mean in your context? I've never seen this.

Re: A case study in PDF forensics: The Epstein PDFs

#133
post #39

Earlier quoted context omitted.

He met with moot ("he is sensitive, be gentile", search on jmail), and within a few days the /pol/ board got created, starting a culture war in the US, leading to Trump getting elected president. Absolutely nuts.

be gentile We're just not going to talk about that one I suppose?

'Sensitive' in this context can mean antisemitic. At least that's how I've heard this joke used.

Re: A case study in PDF forensics: The Epstein PDFs

#134

Earlier quoted context omitted.

I think some of the released documents included images of victims, which where redacted. So it's not necessarily malicious removals

If we're assuming they didn't leave victims unredacted on purpose

Pretty devious tactic if so. Chilling effect on both any further witnesses and anybody interested in archiving the data (gives them an ethical conundrum at least). In addition to giving them (the feds) a convenient excuse to take down random docs.

Re: A case study in PDF forensics: The Epstein PDFs

#135
post #97

Earlier quoted context omitted.

Depending on their technical capability, yes. I mean even in this thread you got what are essentially one-liners to do it. Definitely less hassle then doing it irl

I know I'm not the brightest bulb by any measure, but do some people really take less than at least a few minutes to come up with one-liners for problems as novel as graphical transformations to PDFs? Maybe if the presumed techie hacker / federal worker took it as an amusing challenge I could see this being done, but genuinely out of pure laziness? That's incredible if true.

It's not a novel problem. But yes, I don't think people quite appreciate how quick and easy it is for people who are in the habit of brewing up one-liners to solve simple problems to do that. I've done it here on HN for jq toy problems before, and I don't really doubt there are people similarly familiar with imagemagick.

Re: A case study in PDF forensics: The Epstein PDFs

#137
post #85
post #44

What is the legal basis for releasing the someone's private files and communications? If they can do it to Epstein, they can do it to you, to the Washington Post journalist, to former President Clinton, etc. Is the scope at least limited somehow? Generally I favor transparency, but of course probably the most important parts are withheld.

He was a pedophile sex trafficker. Epstein and his clients deserve zero privacy.

You’ve sidestepped the important part of the question.

Re: A case study in PDF forensics: The Epstein PDFs

#138

Earlier quoted context omitted.

be gentile We're just not going to talk about that one I suppose?

'Sensitive' in this context can mean antisemitic. At least that's how I've heard this joke used.

Must be russian or qatari humour <|:o)

Re: A case study in PDF forensics: The Epstein PDFs

#139
post #58

Earlier quoted context omitted.

People always claimed this as a data leak vector but I've always been sceptical. Like just writing style and vocabulary is probably extremely shared among too many people to narrow it down much. (How people that you know could have written this reply?) The counter argument is that he had a very specific style in his mail so maybe this is a special case.

If you have a large enough set to test against and a specific person you are looking for, this is totally doable currently.

Not just a test set, but enough of a set to search through and compare against. Several pages of in-depth writing isn't anywhere near sufficient, even when limiting the search space to ~10k people.

Re: A case study in PDF forensics: The Epstein PDFs

#140
post #72
post #58

Earlier quoted context omitted.

People always claimed this as a data leak vector but I've always been sceptical. Like just writing style and vocabulary is probably extremely shared among too many people to narrow it down much. (How people that you know could have written this reply?) The counter argument is that he had a very specific style in his mail so maybe this is a special case.

this is a well-studied field (stylometry). when combining writing styles, vocabulary, posting times, etc. you absolutely can narrow it down to specific people. even when people deliberately try to feign some aspects (e.g. switching writing styles for different pseudonyms), they will almost always slip up and revert to their most comfortable style over time. which is great, because if they aren't also regularly changi…

Stylometry is okay if you're trying to deanonymize a large enough sample text. A reddit account would be doable. But individual 4chan posts? You barely have enough content within the text limit.
Post reply on HN