Live data from Hacker News

A case study in PDF forensics: The Epstein PDFs

pdfa.org

241–250 of 254 posts

Re: A case study in PDF forensics: The Epstein PDFs

#241
post #87

Earlier quoted context omitted.

Not that I can speak from personal experience or anything... But somebody on an email chain may have requested a scanned version of the document to ensure there is no metadata and the employee might have found it easier to just flatten the pdf and apply a graphical filter to make the document appear like a scanned document. There might even be a webtool available somewhere to do so, I wouldn't know...

> the employee might have found it easier to just flatten the pdf and apply a graphical filter to make the document appear like a scanned document Is that remotely plausible? I can't imaging faking a scan being easier than just walking down the hall to the copier room.

Look, what I'm saying is that I don't have a scanner at home or at work and I've find this.

Re: A case study in PDF forensics: The Epstein PDFs

#242
post #144

Has anyone analysed JE's writing style and looked for matches in archived 4chan posts or content from similar platforms? Same with Ghislaine, there should be enough data to identify them atp right? I don't buy the MaxwellHill claims for various reasons but it doesn't mean there's nothing to find.

There was a post on here about a project in stylometry that analyzed HN users comment history. The tool helped find accounts that had an extremely similar writing style to a given account. The site was soon removed due to privacy concerns but many users with multiple account attested to its accuracy https://news.ycombinator.com/item?id=33755016 It turns out stylometry is actually a pretty well-developed field. It mak…

On the one side it's a shame this tool was removed because it's very interesting, but on the other hand, the main use case would likely abuse and (cyber)stalking.

That said, best to assume that the various government agencies have tools like this, and better - if you're trying to hide your identity online, don't just change users or go through VPNS/proxies/TOR but change your writing style too.

(Also I'm convinced most VPNs/ proxies / TOR nodes / public access points are honeypots)

Re: A case study in PDF forensics: The Epstein PDFs

#243

Earlier quoted context omitted.

If you have a large enough set to test against and a specific person you are looking for, this is totally doable currently.

Of course it's doable. The question is how reliable the results are.

Not to mention, I'd argue that most people have a (subtly?) different writing style depending on where they post and to who they talk to.

Re: A case study in PDF forensics: The Epstein PDFs

#244

Earlier quoted context omitted.

Stylometry is extremely sophisticated even with simple n-gram analysis. There's a demo of this that can easily pick out who you are on HN just based on a few paragraphs of your own writing, based on N-gram analysis. https://news.ycombinator.com/item?id=33755016 You can also unironically spot most types of AI writing this way. The approaches based on training another transformer to spot "AI generated" content are wron…

> You can also unironically spot most types of AI writing this way. I have no idea if specialized tools can reliably detect AI writing but, as someone whose writing on forums like HN has been accused a couple of times of being AI, I can say that humans aren't very good at it. So far, my limited experience with being falsely accused is it seems to partly just be a bias against being a decent writer with a good vocabul…

Thing is, people are on the lookout for obvious AI and I'm sure they have been successful a few times. But this is like confirmation bias, they will never know whether they saw / read something AI generated if they didn't clock it in the first place.

I'm on Reddit too much and a few times there were memes or whatever that were later on pointed out to be AI. And that's the ones that had tells, more and more (and as price goes down / effort/expenditure increases) it will become harder to impossible to tell.

And I have mixed feelings. I don't mind so much for memes, there's little difference between low-effort image editing and low-effort image generation IMO. There's the "advice" / "story" posts which for a long time now have been more of a creative writing effort than true stories, it's a race to the bottom already and AI will only try and accellerate it. But sometimes it's entertaining.

But "fake news" is the dangerous one, and I'm disappointed that combating this seemed to be a passing fad now that the big tech companies and their leaders / shareholders have bent the knee to regimes that are very interested in spreading disinformation/propaganda to push their agenda under people's skins subtly. I'm surprised it's not more egregious tbh, but maybe it's because my internet bubbles are aligned with my own opinions/morals/etc at the moment.

Re: A case study in PDF forensics: The Epstein PDFs

#245
post #228

Earlier quoted context omitted.

Im not too deep into USA politics and have very very bad memory so i dont remember how it went down. The Wikipedia article you linked says it was signed by trump. >and then withheld files. So did he sign that willingly in the end? Did he have to sign it? Did he cave because he said publicly he would?

A super majority in Congress voted for the files to be the released, which is enough to override a veto, so he had to sign it to save face.

thx for the info!!

Re: A case study in PDF forensics: The Epstein PDFs

#246

Has anyone analysed JE's writing style and looked for matches in archived 4chan posts or content from similar platforms? Same with Ghislaine, there should be enough data to identify them atp right? I don't buy the MaxwellHill claims for various reasons but it doesn't mean there's nothing to find.

> I don't buy the MaxwellHill claims for various reasons

Why not? Clear motive, matching timeline, mentions of that reddit account in the released FBI documents of her case

Re: A case study in PDF forensics: The Epstein PDFs

#247
post #184

Earlier quoted context omitted.

The old trick years ago was to translate from English to different language and back (possibly repeating). I'd be curious how helpful it is against stylometry detection? If you want to be grouped with foreigners who don't know English, it might work well, although word choices may still be distinctive enough to differentiate even when translated.

Assuming the source language is English, going to a romance language and back wouldn't be too hard grammar wise, but could easily wipe out a lot of non-Latin-descended words if you use the right approach to translation.

"We are pretty sure the target is an Anglishman, which definitely narrows our field but maybe a bit too much."

Re: A case study in PDF forensics: The Epstein PDFs

#248

Has anyone analysed JE's writing style and looked for matches in archived 4chan posts or content from similar platforms? Same with Ghislaine, there should be enough data to identify them atp right? I don't buy the MaxwellHill claims for various reasons but it doesn't mean there's nothing to find.

> I don't buy the MaxwellHill claims for various reasons Why not? Clear motive, matching timeline, mentions of that reddit account in the released FBI documents of her case

I was there for the original thread making the connection so I got a very fresh look at the profile. The user was consistently referring to being in dental school. A lot of posting, and not in ways that would influence opinions. Maybe a cover for more secretive mod actions, but it'd be a wastefully excessive cover.

Other mods knew them personally and were still in contact. The user claims they heard of the rumor and decided not to reactivate for the lulz.

I am not familiar with the mod side of reddit - couldn't fellow mods audit her mod action logs to find more juicy details we would have heard about by now?

If Maxwell is indeed a spy and doing what she is claimed to do, it is highly unlikely that she'd put her last name and a reference to her specific family's property in her username. This would be a glaringly arrogant choice for someone who had been groomed from an early age for spycraft, and who had any degree of oversight.

If she were part of a spy network, they would be highly remiss not to commandeer the account at the time of her arrest to avoid suspicion unless they were completely incompetent.

I am mostly familiar with cold war espionage so it just doesn't sound like the general MO to me. Unless Opsec or whatever has badly decayed since then. That's not impossible.

The mentions of the account in the files are from anonymous tips, some of which are highly absurd. They vetted a lot of tips, and I saw no information in the new releases indicating they thought it held water. We've seen the subpoena and IP tracking for the Epstein prison guard whistleblower, but no such thing on this topic.

Re: A case study in PDF forensics: The Epstein PDFs

#249

Has anyone analysed JE's writing style and looked for matches in archived 4chan posts or content from similar platforms? Same with Ghislaine, there should be enough data to identify them atp right? I don't buy the MaxwellHill claims for various reasons but it doesn't mean there's nothing to find.

The writing style is rather interesting. Epstein seems borderline dyslexic, but almost none of the emails I've seen are written in a coherent way, regardless of the sender. Either people on that level rarely write anything on their own and have completely forgotten how to construct proper sentences or maybe that just how they communicate. Sort of language internal to the group.

I had a boss who was too impatient to hear a full sentence most of the time, and respected absolutely no one. She typed like this.

Re: A case study in PDF forensics: The Epstein PDFs

#250

Earlier quoted context omitted.

> You can also unironically spot most types of AI writing this way. I have no idea if specialized tools can reliably detect AI writing but, as someone whose writing on forums like HN has been accused a couple of times of being AI, I can say that humans aren't very good at it. So far, my limited experience with being falsely accused is it seems to partly just be a bias against being a decent writer with a good vocabul…

> I can say that humans aren't very good at it You're assuming the people making accusations of posts being written by AI are from humans (which I agree are not good at making this determination). However, computers analyzing massive datasets are likely to be much better at it , and this can also be a Werewolf/Mafia/Killers-type situation where AI frequently accuses posters it believes are human, of being AI, to dimi…

Are you impugning intent on the LLM's part?
Post reply on HN