Live data from Hacker News

UK scientists ask for help in transcribing 200 years of rainfall data

bbc.com

41–43 of 43 posts

Re: UK scientists ask for help in transcribing 200 years of rainfall data

#41

Earlier quoted context omitted.

What accuracy do you think humans have? If I give you 1000 handwritten numbers, do you think you'll make less than 10 mistakes?

The forms are in open access, anyone can fill one form, so they surely did consider variations in accuracy (if not plain sabotage). They are probably submitting the same data to multiple people and the cross-checking the submissions. Also, the totals help to filter out submissions with mistakes. One can also add bound checking for individual values (misplaced decimal point, extra digits etc.). I would love to see how…

This is exactly how zooniverse works. The first thing you do is go through a few known pieces of data to verify you know how to identify things correctly (more important for galaxy types, etc).

As you start labeling data everything has to reach a consensus. I'm not sure exactly how it works, but it does have multiple people verify each piece of data.

Re: UK scientists ask for help in transcribing 200 years of rainfall data

#42
post #23

Why science must always be done for free, and all the other non scientific stuff, including things dificult to explain logically, are gladly and generously overpaid? Universities take an economic profit from this volunteer work. They can just share some and hire a few students to do it. And if there is a scarcity of more volunteers, maybe they should think about if repeatedly firing, whimsically dismantling expert te…

How about the transcriber getting free tuition, and a citation in the published data.

Re: UK scientists ask for help in transcribing 200 years of rainfall data

#43
post #31
post #13

Earlier quoted context omitted.

Is there a technology that would help ocr-ing standardized forms? I’ve got a couple thousand pages of historical train schedules that i would like to digitize (tabulated, printed data, including symbols/icons), but I’m not sure how to automatically recognize the structured data.

I'm curious, what's your interest in the train schedules? One day at an antique shop, I came across a book from ~1910 which had hundreds of pages of annual reports from railroads with many metrics we'd expect to see in the 10-K reports public companies file. The book was published annually, but had much of its data in tables with grouped headers and cells, which could make automated OCR-ing with a good (useful) end r…

I’ve got European rail schedules, which gives an indication of historical travel patterns. Specifically looking into the development of overnight travel.
Post reply on HN