Live data from Hacker News

Can you read this cursive handwriting? The National Archives wants your help

smithsonianmag.com

101–110 of 267 posts

Re: Can you read this cursive handwriting? The National Archives wants your help

#101
post #39

Before commenting asking about why they don't just use LLMs, please note that the article specifically calls out that they do, but it's not always a viable solution: > The agency uses artificial intelligence and a technology known as optical character recognition to extract text from historical documents. But these methods don’t always work, and they aren’t always accurate. The document at the top is likely an especi…

OK, fair enough, but can you find one in this article that's hard for an LLM? The gnarliest one I saw, 4o handled instantly, and I went back and looked carefully at the image and the text and I'm sold. Like if this is a crowdsourcing project, why not do a first pass with an LLM and present users with both the image and the best-effort LLM pass? Later I signed up, went to the current missions, and they all seem to pos…

One that require additional work beyond simply feeding the image into the model would be this example which is a mix of barely legible hand written cursive and easy to read typed form. [0] Initially 4o just transcribes (successfully) the bottom half of the text and has to be prompted to attempt the top half at which point it seems to at best summarize the text instead of giving a direct transcription. [1] In fact it seems to mix up some portions of the latter half of the typed text with the written text in the portion of it's "transcription" about "reduced and indigent circumstances".

[0] https://catalog.archives.gov/id/54921817?objectPage=8&object...

[1] Reproducing here since I cannot share the chat since it has user uploaded images. " The text in the top half of the image is handwritten and partially difficult to read due to its cursive style and some smudging. Here's my best transcription attempt for the top section:

...resident within four? years, swears and says that the name of the John Hopper mentioned in the foregoing declaration is the same person, and he verily believes the facts as stated in the declaration are true.

He further swears that the said John Hopper is in reduced and indigent circumstances and requires the aid of his country.

The declarant further swears he has no evidence now in his power of service, except the statement of Capt. (illegible name), as to his reduced circumstances ...

Sworn to before me, this day...

Some parts remain unclear due to the handwriting, but let me know if you'd like me to attempt further clarification on specific sections!"

Re: Can you read this cursive handwriting? The National Archives wants your help

#102

Before commenting asking about why they don't just use LLMs, please note that the article specifically calls out that they do, but it's not always a viable solution: > The agency uses artificial intelligence and a technology known as optical character recognition to extract text from historical documents. But these methods don’t always work, and they aren’t always accurate. The document at the top is likely an especi…

Something about extraordinary claims and extraordinary evidence? The evidence presented, a seemingly easily transcribed image, is hardly persuasive.

Some are significantly harder to read. I took the page below and tried to get GPT 4o to transcribe it and it basically couldn't do it. I'm not going to sit and prompt hack for ages to see if it can but it seems unable to tackle the handwritten text at the top. When I first just fed it the image and asked for a transcription it only (but successfully) read the bottom portion, prompted for a transcription of the top it dropped into more of a summary of the whole document mainly pulling some phrases from the bottom text. (Sadly can't share it but I copied it's reply out in a comment upthread) [0]

It was more successful at a few others I tried but it's still a task that requires manual processing like a lot of LLM output to check for accuracy and prompt modification to get it to output what you need for some documents.

https://catalog.archives.gov/id/54921817?objectPage=8&object...

[0] https://news.ycombinator.com/item?id=42746490

Re: Can you read this cursive handwriting? The National Archives wants your help

#103
post #39

Before commenting asking about why they don't just use LLMs, please note that the article specifically calls out that they do, but it's not always a viable solution: > The agency uses artificial intelligence and a technology known as optical character recognition to extract text from historical documents. But these methods don’t always work, and they aren’t always accurate. The document at the top is likely an especi…

OK, fair enough, but can you find one in this article that's hard for an LLM? The gnarliest one I saw, 4o handled instantly, and I went back and looked carefully at the image and the text and I'm sold. Like if this is a crowdsourcing project, why not do a first pass with an LLM and present users with both the image and the best-effort LLM pass? Later I signed up, went to the current missions, and they all seem to pos…

> Like if this is a crowdsourcing project, why not do a first pass with an LLM and present users with both the image and the best-effort LLM pass?

Possibly for the reason that came up in your other post: you mentioned that you spot checked the result.

Back when I was in historical research, and occasionally involved in transcription projects, the standard was 2-3 independent transcriptions per document.

Maybe the National Archive will pass documents to an LLM and use the output as 1 of their 2-3 transcriptions. It could reduce how many duplicate transcriptions are done by humans. But I'll be surprised if they jump to accepting spot checked LLM output anytime soon.

Re: Can you read this cursive handwriting? The National Archives wants your help

#104
post #32

Earlier quoted context omitted.

I would challenge you to find a picture of text that you think a human can read and OCR cannot. I’m happy to demonstrate. The text shown in this article is trivial.

> I would challenge you to find a picture of text that you think a human can read and OCR cannot. Are you aware of CAPTCHA[0] images? 0 - https://en.wikipedia.org/wiki/CAPTCHA

Text that is _intentionally constructed_ to fool computers but not humans is obviously out of scope. But they’re generally easily solved with OCR these days anyway.

Re: Can you read this cursive handwriting? The National Archives wants your help

#105
post #32

Earlier quoted context omitted.

I would challenge you to find a picture of text that you think a human can read and OCR cannot. I’m happy to demonstrate. The text shown in this article is trivial.

The archivists themselves say that they run into such texts often enough that this program was needed: > The agency uses artificial intelligence and a technology known as optical character recognition to extract text from historical documents. But these methods don’t always work, and they aren’t always accurate. They are absolutely aware of the advances in these tools, so if they say they're not completely there yet…

Then please provide a single example that we can’t instantly solve. Happy to prove them wrong.

Re: Can you read this cursive handwriting? The National Archives wants your help

#106
post #29

Earlier quoted context omitted.

There are conceivable reasons why they may be telling a half truth here. Just engaging the public is a worthy goal here.

> There are conceivable reasons why they may be telling a half truth here. Just engaging the public is a worthy goal here. Asserting an ulterior motive without supporting proof is to engage in conspiracy theories. Sometimes a cigar is just a cigar.[0] 0 - https://quoteinvestigator.com/2011/08/12/just-a-cigar/

The alternative is me saying that appealing to their “expertise” is an appeal to authority fallacy that flies in the face of general evidence that modern OCR is far better than humans at character recognition. Especially random non specialized humans.

Re: Can you read this cursive handwriting? The National Archives wants your help

#107
post #5

I don’t think I believe that OCR can’t do it but random humans can OCR is VERY good

> I don’t think I believe that OCR can’t do it but random humans can Considering the people involved are experts in their field, are certainly aware of OCR capabilities, and have publicized a need thusly: ... the National Archives is looking for volunteers who can help transcribe and organize its many handwritten records ... Perhaps "random humans" can perform tasks which could reshape your belief: > OCR is VERY good

Also, you seem to have taken issue with the phrase “random humans” because you’re confused at what’s being done here. It is random humans. Non experts.

Experts are asking for the help of non experts.

> Anyone with an internet connection can volunteer to transcribe historical documents and help make the archives’ digital catalog more accessible

Re: Can you read this cursive handwriting? The National Archives wants your help

#108
post #40
post #5

I don’t think I believe that OCR can’t do it but random humans can OCR is VERY good

I've been trying every state of the art OCR solution on my students' handwritten essays for fifteen years and have yet to find anything even close to acceptable.

What methods have you tried?

Re: Can you read this cursive handwriting? The National Archives wants your help

#109
To tptacek and other guys who seem to have unwavering trust in OCRs/LLMs, as well as to opposite party who think that technology is not there yet — you are all partially right, but somehow fail to hear each other while also spending time on baseless arguing instead of factual examples and attempts to find common truth.

Can it be used to greatly simplify efforts by getting through boilerplate? — Yes.

Should the result be reviewed and proof-read by human? — Also yes.

---

Here subtle one: https://catalog.archives.gov/id/34384201?objectPage=40

Here is (one of) transcripts made by `o1-pro`:

  (2)

  …and I don’t know whether it can be reset for a
  date in December or not. Cornell seemed
  anxious that it should not come up too close to Christmas,
  and of course new suspicion [would be aroused?] [about?] him.
  I will take this up with the Judge as soon as I can get rid of the brief.
  Meanwhile I would like to know whether there is anything else
  in which I can be useful to you, since it behooves me
  in ways of uncomfortable relations with the present management.

  Are you going East in December?
  Has any word come from Hagerman?
  Were there any noteworthy developments at the hearings
  on the [Teapot?] trial?

  I have no inclination yet whether Wheeler will be wanted in
  Washington, but the chances are that he will not.

  With regards to all the brethren and [flock?], I am

  very sincerely yours,
  George A. H. Fraser
I'm not native english speaker, but even I can read where it is wrong. I'll leave it to be an excercise for the reader to find out mistakes, but it is certainly not a Teapot trial.

Somehow GPT-4o performs better on this example and fails only on "New Mexican practise" part.

Post reply on HN