Live data from Hacker News

Recreating Epstein PDFs from raw encoded attachments

neosmart.net

51–60 of 224 posts

Re: Recreating Epstein PDFs from raw encoded attachments

#51
post #30
post #20

Earlier quoted context omitted.

I mean, the internet is finding all her mistakes for her. She is actually doing alright with this. Crowdsource everything, fix the mistakes. lol.

This would be funnier if it wasn’t child porn being unredacted by our government

[flagged]

Re: Recreating Epstein PDFs from raw encoded attachments

#52

On one hand, the DOJ gets shit because it was taking too long to produce the documents, and then on another, they get shit because there are mistakes in the redacting because there are 3 million pages of documents.

What they are redacting is pretty questionable though. Entire pages being suspiciously redacted with no explanation (which they are supposed to provide). This is just my opinion, but I think it's pretty hard to defend them as making an honest and best effort here. Remember they all lied about and changed their story on the Epstein "files" several times now (by all I mean Bondi, Patel, Bongino, and Trump).

It's really really hard to give them the benefit of the doubt at this point.

Re: Recreating Epstein PDFs from raw encoded attachments

#54
post #13

> …but good luck getting that to work once you get to the flate-compressed sections of the PDF. A dynamic programming type approach might still be helpful. One version or other of the character might produce invalid flate data while the other is valid, or might give an implausible result.

Time to flex those Leetcode skills.

Re: Recreating Epstein PDFs from raw encoded attachments

#56
post #33

Given how much of a hot mess PDFs are in general, it seems like it would behoove the government to just develop a new, actually safe format to standardize around for government releases and make it open source. Unlike every other PDF format that has been attempted, the federal government doesn't have to worry about adoption.

JPEG?

Lossy

Re: Recreating Epstein PDFs from raw encoded attachments

#57

Nerdsnipe confirmed :) Claude Opus came up with this script: https://pastebin.com/ntE50PkZ It produces a somewhat-readable PDF (first page at least) with this text output: https://pastebin.com/SADsJZHd (I used the cleaned output at https://pastebin.com/UXRAJdKJ mentioned in a comment by Joe on the blog page)

So it was a public event attended by 450 people:

https://www.mountsinai.org/about/newsroom/2012/dubin-breast-...

https://www.businessinsider.com/dubin-breast-center-benefit-...

Even names match up, but oddly the date is different.

Re: Recreating Epstein PDFs from raw encoded attachments

#58
post #43

Earlier quoted context omitted.

The US administration is, at present, regularly violating the law and ignoring court orders. Indeed, these very releases are patently in violation of multiple federal laws -- they're simultaneously insufficiently-responsive to meet the requirements of the law requiring the release of the files and fall afoul of CSAM laws by being incompletely redacted. The challenge, as we're all experiencing together, is that the la…

Can you provide a couple examples of the laws they're violating?

How about court orders?

https://www.cbsnews.com/minnesota/news/ice-violations-judge-...

> ICE has likely violated more court orders in January 2026 than some federal agencies have violated in their entire existence," Schiltz said, adding that he counted 96 court orders that ICE has violated in 74 cases.

https://www.cbsnews.com/news/frustrations-from-judge-prosecu...

Re: Recreating Epstein PDFs from raw encoded attachments

#59

This is one of those things that seems like a nerd snipe but would be more easily accomplished through brute forcing it. Just get 76 people to manually type out one page each, you'd be done before the blog post was written.

Or one person types 76 pages. This is a thing people used to do, not all that infrequently. Or maybe you have one friend who will help–cool, you just cut the time in half.

Typing 76 pages is easy when it's words in a language you understand. WPM is going to be incredibly slow when you actually have to read every character. On top of that, no spaces and no spellcheck so hopefully you didn't miss a character.
Post reply on HN