Live data from Hacker News

Ask HN: OCR for 100 year old (German) handwritten cursive script?

news.ycombinator.com

41–45 of 45 posts

Re: Ask HN: OCR for 100 year old (German) handwritten cursive script?

#42

You could try something like https://aws.amazon.com/textract/ or https://cloud.google.com/vision/docs/handwriting . Both have support for modern handwriting. I don't know if it will work with a script written a century ago though.

if it's https://en.wikipedia.org/wiki/Sütterlin I doubt anything trained on current script would make any more sense of it than we do

There's models for this, see https://readcoop.eu/model/german-kurrent-and-sutterlin-17th-...

Re: Ask HN: OCR for 100 year old (German) handwritten cursive script?

#43
post #21

I've had surprisingly good results with https://readcoop.eu/transkribus/ I was going back in time with a family research until I couldn't identify a single word anymore. The 'AI' could.

My colleagues are mostly using transkribus for handwriting. I work at a library.

Re: Ask HN: OCR for 100 year old (German) handwritten cursive script?

#44
As others have mentioned, Transkribus works pretty well for handwritten text recognition. You can also train your own model if you have enough source material.

If the documents you have are able to be made public, you could upload them to Wikimedia Commons and use https://ocr.wmcloud.org/ — you can use Transkribus via that. (Disclosure: I'm an engineer working on the Wikimedia OCR project.)

Re: Ask HN: OCR for 100 year old (German) handwritten cursive script?

#45

You probably want to put it in front of an actual person and get them to transcribe it for you. I don't think there's any off the shelf OCR that will work particularly well for it. I have a close family member who is a historian and frequently read and transcribed mid 19th to early 20th century German handwriting for his work. Many historians and archivists in Germany would have the ability to transcribe this for you…

This is probably the best option. I can't find it now, but in 2020 on either Slashdot.org or here, there was a project trying to transcribe hundreds or thousands of old British rainfall records.

The researchers made digital scans and posted the images online and had random users around the world transcribe them. They didn't care if a user did one or hundreds. I did about 20-50 before they were finished. What would have taken a paid team years was completed in only a week.

Does anyone know a link to the article that announced it?

Post reply on HN