What a sad state for humanity that we have to resort to this sort of OCR/scrapping instead of the original data being released in a machine readable format in the first place.
1) There's plenty of old data out there. Newspaper scans from the days before computers, or digitalization of the newspaper process. Or the original files simply got lost, so manually scanned pages is all you have.
2) There could be policies about making the data public, but in a way that discourages data scraping.
3) The providers of the data simply don't have the resources or incentives to develop a working API.
And many more.