Live data from Hacker News

UK scientists ask for help in transcribing 200 years of rainfall data

bbc.com

31–40 of 43 posts

Re: UK scientists ask for help in transcribing 200 years of rainfall data

#31
post #13

Why hasn't this been automated? Digit recognition technology is pretty good and it seems like the forms are pretty standardized.

Is there a technology that would help ocr-ing standardized forms? I’ve got a couple thousand pages of historical train schedules that i would like to digitize (tabulated, printed data, including symbols/icons), but I’m not sure how to automatically recognize the structured data.

I'm curious, what's your interest in the train schedules?

One day at an antique shop, I came across a book from ~1910 which had hundreds of pages of annual reports from railroads with many metrics we'd expect to see in the 10-K reports public companies file.

The book was published annually, but had much of its data in tables with grouped headers and cells, which could make automated OCR-ing with a good (useful) end result challenging.

I think it'd be interesting to map out the Railroad consolidation, track all their financial metrics over time, and do some level of forensic accounting to see if/which companies probably had funny business going on.

Re: UK scientists ask for help in transcribing 200 years of rainfall data

#32
post #23

Why science must always be done for free, and all the other non scientific stuff, including things dificult to explain logically, are gladly and generously overpaid? Universities take an economic profit from this volunteer work. They can just share some and hire a few students to do it. And if there is a scarcity of more volunteers, maybe they should think about if repeatedly firing, whimsically dismantling expert te…

> Universities take an economic profit from this volunteer work

This is untrue in many places. Like, not even close. If anything, everyone else economically benefits off the back of groundbreaking and blue-skies research.

Re: UK scientists ask for help in transcribing 200 years of rainfall data

#34
post #23

Why science must always be done for free, and all the other non scientific stuff, including things dificult to explain logically, are gladly and generously overpaid? Universities take an economic profit from this volunteer work. They can just share some and hire a few students to do it. And if there is a scarcity of more volunteers, maybe they should think about if repeatedly firing, whimsically dismantling expert te…

> Universities take an economic profit from this volunteer work This is untrue in many places. Like, not even close. If anything, everyone else economically benefits off the back of groundbreaking and blue-skies research.

For efforts like this, the University really only gets some news recognition blips that people quickly forget.

The PI may have a slightly sexier project if they successfully pull off said endeavor which leads to increased competitiveness in review processes, increased likelihood of funding, and then the University likely gets some overhead in many cases.

The real question for me isn't from Universities making money (which is highly indirect and likely fairly insignificant in volume). The real question is why research is so frequently underfunded and a low priority in this country. There can be more budget for blueskies research but ROI is long, unknown, and research is treated like a business. Research should be treated like an endeavor for knowledge where people doing it need to eat, have a home, etc. and realize it's being done and released to the public for everyone.

Re: UK scientists ask for help in transcribing 200 years of rainfall data

#36
10+ years ago, Ancestry.com hired people in Asia to transcribe U.S. census returns and vital records. They also used volunteers: http://blogs.ancestry.com/circle/?p=2321

Not sure how it's done now, but I suspect it's very hard to automate owing to inconsistent handwriting styles, unusual names for people and places, footnotes and abbreviations, etc.

Re: UK scientists ask for help in transcribing 200 years of rainfall data

#37
post #23

Why science must always be done for free, and all the other non scientific stuff, including things dificult to explain logically, are gladly and generously overpaid? Universities take an economic profit from this volunteer work. They can just share some and hire a few students to do it. And if there is a scarcity of more volunteers, maybe they should think about if repeatedly firing, whimsically dismantling expert te…

Zooniverse was a direct response to a PhD project that Chris Lintott was supervising. His student was tasked with classifying galaxies and in the end decided that the volume of data was just too great. There is a good supply of volunteers, and the tasks are enjoyed by those who undertake them. Where volunteers identify notable things, they are also credited by name in the resulting papers. > maybe they should think a…

> I don't see much evidence for this happening in the UK, where the project is based.

Reduction in a 40% of UK funding applicants in the last years:

https://www.bbc.com/news/science-environment-50044659

Lets see how the nice promises will perform at the end

https://www.sciencemag.org/news/2020/03/uk-cues-big-funding-...

Re: UK scientists ask for help in transcribing 200 years of rainfall data

#38
post #23

Why science must always be done for free, and all the other non scientific stuff, including things dificult to explain logically, are gladly and generously overpaid? Universities take an economic profit from this volunteer work. They can just share some and hire a few students to do it. And if there is a scarcity of more volunteers, maybe they should think about if repeatedly firing, whimsically dismantling expert te…

Zooniverse was a direct response to a PhD project that Chris Lintott was supervising. His student was tasked with classifying galaxies and in the end decided that the volume of data was just too great. There is a good supply of volunteers, and the tasks are enjoyed by those who undertake them. Where volunteers identify notable things, they are also credited by name in the resulting papers. > maybe they should think a…

Ha, funny seeing this mentioned on HN, this student was my Professor for Astrophysics during my time in this field. He was a great mentor and taught me (and others) quite a lot, in particular about outreach and science communication. The Zooniverse project is amazing and citizen science in general is a great idea. However, you mention the crux of the issue already: the tasks need to be enjoyable. Looking at galaxies that potentially no other human ever laid eyes on is pretty much as exciting as it may get. A former colleague of mine (in the group of the mentioned student) worked on this interesting structure https://en.wikipedia.org/wiki/Hanny%27s_Voorwerp that was discovered by one citizen scientist in the Zooniverse project (hence the name). It's one of the coolest success stories of crowd-sourced research I've heard about.

Re: UK scientists ask for help in transcribing 200 years of rainfall data

#39
post #3

Why hasn't this been automated? Digit recognition technology is pretty good and it seems like the forms are pretty standardized.

Probably because 99% accuracy isn't good enough.

Outliers would be easy to spot.

Maybe they should split in lines each figure, make a coordinates map of the entire figure or so, make a copy and applying a bulk search with a machine for crossing the map and annotate as many undoubtely identifiable numbers as possible. Then paint it in a different color easy to filter or hide it and remember its position.

And then add humans to focuse only in the remaining dificult cases and outliers armed with a reference sample chart. This way an human would need to focus its eyes in 20 characters/image (instead 60 or 100). They would accumulate more completed figures faster and obtain a bigger sense of reward. If the human can't recognize the number, could put a ? in the chain and move on.

Thus the machine would annotate for example 1_3, 4_67_, _8 and the human would write: 5?01 for the same line.

Just an idea, don't know if designing that is easy-peasy or really defiant but the number of possible values is limited in any case, so should be possible to train a machine to recognize some single characters or even entire numbers. Specially if written for the same people.

Re: UK scientists ask for help in transcribing 200 years of rainfall data

#40
Some kind of input validation and user feedback would be welcome. Show a green check mark, if the input annual sum equals the computed sum of the input monthly sums. Typically, all values have two decimal places, thus, validate that there are no more and no less in the user input.
Post reply on HN