Why hasn't this been automated? Digit recognition technology is pretty good and it seems like the forms are pretty standardized.
Is there a technology that would help ocr-ing standardized forms? I’ve got a couple thousand pages of historical train schedules that i would like to digitize (tabulated, printed data, including symbols/icons), but I’m not sure how to automatically recognize the structured data.
One day at an antique shop, I came across a book from ~1910 which had hundreds of pages of annual reports from railroads with many metrics we'd expect to see in the 10-K reports public companies file.
The book was published annually, but had much of its data in tables with grouped headers and cells, which could make automated OCR-ing with a good (useful) end result challenging.
I think it'd be interesting to map out the Railroad consolidation, track all their financial metrics over time, and do some level of forensic accounting to see if/which companies probably had funny business going on.