Ask HN: OCR for 100 year old (German) handwritten cursive script?
21–30 of 45 posts
Re: Ask HN: OCR for 100 year old (German) handwritten cursive script?
#22You can train Tesseract to recognize Handwriting[1], but the first and most important step would be the preprocessing of your documents. I would recommend to start with a local adaptive thresholding algorithm[2] like Sauvola for binarization. The preprocessing steps would be[3] 1) Binarization 2) Skew Correction 3) Noise Removal 4) Thinning and Skeletonization Probably you are facing "Sütterlin"[4], which differs qui…
Re: Ask HN: OCR for 100 year old (German) handwritten cursive script?
#23General purpose open-source OCR solutions like Tesseract, TrOCR, etc will probably not be as good as the cloud ones, based on my experience.
There's some specialized research work out there for antique manuscripts, but that will require some digging on your part with an uncertain outcome. I think at that point, I would also look into manual transcription - for 200 pages, it might be reasonably affordable.
Re: Ask HN: OCR for 100 year old (German) handwritten cursive script?
#24GPT-4 Vision. I have seen some examples of middly agy looking pages tried.
Re: Ask HN: OCR for 100 year old (German) handwritten cursive script?
#25I don't know if it will be significantly different than what Google Translate does, but I would give the major cloud vendors (Google, Amazon, Microsoft and I guess OpenAI/ChatGPT) OCR services a shot. It's pretty simple and cheap to do (like, about a dollar for the whole thing). Last time I compared them, Google's OCR came out ahead, but it's task-dependent so in your case it might be different. General purpose open-…
Re: Ask HN: OCR for 100 year old (German) handwritten cursive script?
#26You can train Tesseract to recognize Handwriting[1], but the first and most important step would be the preprocessing of your documents. I would recommend to start with a local adaptive thresholding algorithm[2] like Sauvola for binarization. The preprocessing steps would be[3] 1) Binarization 2) Skew Correction 3) Noise Removal 4) Thinning and Skeletonization Probably you are facing "Sütterlin"[4], which differs qui…
Re: Ask HN: OCR for 100 year old (German) handwritten cursive script?
#27Does it look like Sütterlin? Are you familiar with it? [] https://en.wikipedia.org/wiki/S%C3%BCtterlin
Not familiar with it, but it doesn’t look like that. I wish haha
Re: Ask HN: OCR for 100 year old (German) handwritten cursive script?
#28They had a contract to index historical French archives composed of handwritten latin documents in elasticsearch.
Depending of the historical relevance of your documents (read: some academic funds), they may be able to help. Doesn't hurt to contact them: