Before commenting asking about why they don't just use LLMs, please note that the article specifically calls out that they do, but it's not always a viable solution: > The agency uses artificial intelligence and a technology known as optical character recognition to extract text from historical documents. But these methods don’t always work, and they aren’t always accurate. The document at the top is likely an especi…
OK, fair enough, but can you find one in this article that's hard for an LLM? The gnarliest one I saw, 4o handled instantly, and I went back and looked carefully at the image and the text and I'm sold. Like if this is a crowdsourcing project, why not do a first pass with an LLM and present users with both the image and the best-effort LLM pass? Later I signed up, went to the current missions, and they all seem to pos…
[0] https://catalog.archives.gov/id/54921817?objectPage=8&object...
[1] Reproducing here since I cannot share the chat since it has user uploaded images. " The text in the top half of the image is handwritten and partially difficult to read due to its cursive style and some smudging. Here's my best transcription attempt for the top section:
...resident within four? years, swears and says that the name of the John Hopper mentioned in the foregoing declaration is the same person, and he verily believes the facts as stated in the declaration are true.
He further swears that the said John Hopper is in reduced and indigent circumstances and requires the aid of his country.
The declarant further swears he has no evidence now in his power of service, except the statement of Capt. (illegible name), as to his reduced circumstances ...
Sworn to before me, this day...
Some parts remain unclear due to the handwriting, but let me know if you'd like me to attempt further clarification on specific sections!"