Earlier quoted context omitted.
My guess is because it’s the Smithsonian, they’re just not willing to trust an LLM’s transcription enough to put their name on it. I imagine they’re rather conservative. And maybe some AI-skeptic protectionist sentiments from the professional archivists. Seems like it could change with time though.
> My guess is because it’s the Smithsonian, they’re just not willing to trust an LLM’s transcription enough to put their name on it. I imagine they’re rather conservative I expect thats a common theme from companies like that, yet I don't think they understand the issue they think they have there. Why not have the LLMs do as much work as possible and have humans review and put their own name on it? Do you think they…
Can you read this cursive handwriting? The National Archives wants your help
121–130 of 267 posts
Re: Can you read this cursive handwriting? The National Archives wants your help
#122Before commenting asking about why they don't just use LLMs, please note that the article specifically calls out that they do, but it's not always a viable solution: > The agency uses artificial intelligence and a technology known as optical character recognition to extract text from historical documents. But these methods don’t always work, and they aren’t always accurate. The document at the top is likely an especi…
OK, fair enough, but can you find one in this article that's hard for an LLM? The gnarliest one I saw, 4o handled instantly, and I went back and looked carefully at the image and the text and I'm sold. Like if this is a crowdsourcing project, why not do a first pass with an LLM and present users with both the image and the best-effort LLM pass? Later I signed up, went to the current missions, and they all seem to pos…
Re: Can you read this cursive handwriting? The National Archives wants your help
#123Before commenting asking about why they don't just use LLMs, please note that the article specifically calls out that they do, but it's not always a viable solution: > The agency uses artificial intelligence and a technology known as optical character recognition to extract text from historical documents. But these methods don’t always work, and they aren’t always accurate. The document at the top is likely an especi…
OK, fair enough, but can you find one in this article that's hard for an LLM? The gnarliest one I saw, 4o handled instantly, and I went back and looked carefully at the image and the text and I'm sold. Like if this is a crowdsourcing project, why not do a first pass with an LLM and present users with both the image and the best-effort LLM pass? Later I signed up, went to the current missions, and they all seem to pos…
Re: Can you read this cursive handwriting? The National Archives wants your help
#124To tptacek and other guys who seem to have unwavering trust in OCRs/LLMs, as well as to opposite party who think that technology is not there yet — you are all partially right, but somehow fail to hear each other while also spending time on baseless arguing instead of factual examples and attempts to find common truth. Can it be used to greatly simplify efforts by getting through boilerplate? — Yes. Should the result…
Re: Can you read this cursive handwriting? The National Archives wants your help
#125This reminded me of something the historian Megan Marshall wrote in the introduction to her book The Peabody Sisters: Three Women Who Ignited American Romanticism (2005): “I became expert in deciphering the sisters’ handwriting, and that of their ancestors, parents, and friends. Each era and each correspondent presented different challenges. Some hands were sprawling, some spindly, some cramped; t ’s went uncrossed a…
Could these be instances of the long s, “ſ”, easily confused with an f?
Re: Can you read this cursive handwriting? The National Archives wants your help
#126Re: Can you read this cursive handwriting? The National Archives wants your help
#127Earlier quoted context omitted.
OK, fair enough, but can you find one in this article that's hard for an LLM? The gnarliest one I saw, 4o handled instantly, and I went back and looked carefully at the image and the text and I'm sold. Like if this is a crowdsourcing project, why not do a first pass with an LLM and present users with both the image and the best-effort LLM pass? Later I signed up, went to the current missions, and they all seem to pos…
Did you actually check it? Sonnet 3.5 generates text that seems legitimate and generally correct, but misreads important details. LLMs are particularly deceptive because they will be internally consistent - they'll reuse the same incorrect name in both places and will hallucinate information that seems legit, but in fact is just made-up.
Re: Can you read this cursive handwriting? The National Archives wants your help
#128Earlier quoted context omitted.
> My guess is because it’s the Smithsonian, they’re just not willing to trust an LLM’s transcription enough to put their name on it. I imagine they’re rather conservative I expect thats a common theme from companies like that, yet I don't think they understand the issue they think they have there. Why not have the LLMs do as much work as possible and have humans review and put their own name on it? Do you think they…
The incident with the lawyers just highlighted the fundamental problem with LLMs and AI in general. They can't be trusted for anything serious. Worse, they give the apppearence of being correct, which leads humans "checkers" into complacency. Total dumpster fire.
I’m trying to do this for old Latin books at the Embassy of the Free Mind in Amsterdam. So many of the books have never been digitized, let alone OCRd or translated. There is a huge amount of work to be done to make these works accessible.
LLMs won’t make it perfect. But isn’t perfect the enemy of the good? If we make it an ongoing project where the source image material is easily accessible (unlike in a normal published translation, where you just have to trust the translator), then the knowledge and understanding can improve over time.
This approach also has the benefit of training readers not to believe everything they read — but to question it and try to get directly at the source. I think that’s a beautiful outcome.
Re: Can you read this cursive handwriting? The National Archives wants your help
#129Prompt: You are a paleologist specializing in analysis of cursive handwriting; tell me what the following text says: (pasting the picture). Output: The following is the declaration of James Lambert, a soldier of the Revolutionary War in North America. The said James Lambert this day personally appeared in the Probate Court of the County of Dearborn in the state of Indiana and at the November Term of said court (1841)…
Re: Can you read this cursive handwriting? The National Archives wants your help
#130Is that true?! US kids don't learn cursive? How do they write?!