Extracting campaign finance data from gnarly PDFs using deep learning
11–17 of 17 posts
Re: Extracting campaign finance data from gnarly PDFs using deep learning
#12Re: Extracting campaign finance data from gnarly PDFs using deep learning
#13how does the performance compare to a simple ocr reader?
Re: Extracting campaign finance data from gnarly PDFs using deep learning
#14Looks like a very shallow article that hasn't actually solved the meat of the problems. Am I reading it wrong?
Heh. It’s my work, so maybe I can clarify the goals. This is meant to be a proof of concept. I took a week and was able to show that relatively simple deep learning techniques are capable of generalizing over unseen form types with high accuracy. I also showed that tokens-plus-geometry is a viable format, and that hand-crafted feature engineering is still necessary (and still used in SOTA approaches). I also believe…
I think the parents critism was more that the article was a little light compared to thr video. For example, you didn't have screenshots of the scanned pdfs in the article.
Re: Extracting campaign finance data from gnarly PDFs using deep learning
#15My personal bet is that you would have probably gotten better results by simply feeding your data into xgboost or lightgbm. This current fad of focusing on neural networks just seems sorta silly when we have simpler and (often) better performing models. Just look at what sorts of models win kaggle competitions.
Re: Extracting campaign finance data from gnarly PDFs using deep learning
#16Re: Extracting campaign finance data from gnarly PDFs using deep learning
#17Deep learning (and indeed any kind of machine learning) is really not necessary here. I wrote a very simple baseline that estimates the total gross amount as the biggest numerical value present (favouring ones that begin with a dollar sign and/or occur more than once). This achieves 93% accuracy (code is here https://gist.github.com/rlmacsween/166b7c1c1b0c5f466a0fe9b46... ) in comparison to the 90% for the deep learn…