Live data from Hacker News

A new Google model is nearly perfect on automated handwriting recognition

generativehistory.substack.com

201–210 of 328 posts

Re: A new Google model is nearly perfect on automated handwriting recognition

#201

I really hope they have because I’ve also been experimenting with LLMs to automate searching through old archival handwritten documents. I’m interested in the Conquistadors and their extensive accounts of their expeditions, but holy cow reading 16th century handwritten Spanish and translating it at the same time is a nightmare, requiring a ton of expertise and inside field knowledge. It doesn’t help that they were of…

You are right to be skeptical. There are plenty of so called windows(or other) web 'os' clones. There were a couple of these posted on HN actually this very year. Here is one example I google dthat was also on HN : https://news.ycombinator.com/item?id=44088777 This is not an OS as in emulating a kernel in javascript or wasm, this is making a web app that looks like the desktop of an OS. I have seen plenty such projec…

Its always amusing when "an app like windows xp" considered hard or challenging somehow.

Literally the most basic html/css, not sure why it is even included in benchmarks.

Re: A new Google model is nearly perfect on automated handwriting recognition

#202
post #176

I really hope they have because I’ve also been experimenting with LLMs to automate searching through old archival handwritten documents. I’m interested in the Conquistadors and their extensive accounts of their expeditions, but holy cow reading 16th century handwritten Spanish and translating it at the same time is a nightmare, requiring a ton of expertise and inside field knowledge. It doesn’t help that they were of…

I'm surprised people didn't click through to the tweet. https://x.com/chetaslua/status/1977936585522847768 > I asked it for windows web os as everyone asked me for it and the result is mind blowing , it even has python in terminal and we can play games and run code in it And of course > 3D design software, Nintendo emulators No clue what these refer to but to be honest it sounds like they've incrementally improved on…

i think google is going to repeat history with gemini.. as in chatgpt, grok, etc will be like altavista, lycos, etc

Re: A new Google model is nearly perfect on automated handwriting recognition

#203
> In tabulating the “errors” I saw the most astounding result I have ever seen from an LLM, one that made the hair stand up on the back of my neck. Reading through the text, I saw that Gemini had transcribed a line as “To 1 loff Sugar 14 lb 5 oz @ 1/4 0 19 1”. If you look at the actual document, you’ll see that what is actually written on that line is the following: “To 1 loff Sugar 145 @ 1/4 0 19 1”. For those unaware, in the 18th century sugar was sold in a hardened, conical form and Mr. Slitt was a storekeeper buying sugar in bulk to sell. At first glance, this appears to be a hallucinatory error: the model was told to transcribe the text exactly as written but it inserted 14 lb 5 oz which is not in the document.

I read the whole reasoning of the blog author after that, but I still gotta know - how can we tell that this was not a hallucination and/or error? There's a 1/3 chance of an error being correct (either 1 lb 45, 14 lb 5 or 145 lb), so why is the author so sure that this was deliberate?

I feel a good way to test this would be to create an almost identical ledger entry, but in a way so that the correct answer after reasoning (the way the author thinks the model reasoned) has completely different digits.

This way there'd be more confidence that the model itself reasoned and did not make an error.

Re: A new Google model is nearly perfect on automated handwriting recognition

#204

Earlier quoted context omitted.

You are right to be skeptical. There are plenty of so called windows(or other) web 'os' clones. There were a couple of these posted on HN actually this very year. Here is one example I google dthat was also on HN : https://news.ycombinator.com/item?id=44088777 This is not an OS as in emulating a kernel in javascript or wasm, this is making a web app that looks like the desktop of an OS. I have seen plenty such projec…

Its always amusing when "an app like windows xp" considered hard or challenging somehow. Literally the most basic html/css, not sure why it is even included in benchmarks.

Those things are LLMs, with text and language at the core of their capabilities. UIs are, notably, not text.

An LLM being able to build up interfaces that look recognizably like an UI from a real OS? That sure suggests a degree of multimodal understanding.

Re: A new Google model is nearly perfect on automated handwriting recognition

#205
post #64

Earlier quoted context omitted.

Here's a thought experiment: if modern machine learning systems existed in the early 20th century, would they have been able to produce an equivalent to the theory of relativity? How about advance our understanding of the universe? Teach us about flight dynamics and take us into space? Invent the Turing machine, Von Neumann architecture, transistors? If yes, why aren't we seeing glimpses of such genius today? If we'v…

>Here's a thought experiment: if modern machine learning systems existed in the early 20th century, would they have been able to produce an equivalent to the theory of relativity? How about advance our understanding of the universe? Teach us about flight dynamics and take us into space? Invent the Turing machine, Von Neumann architecture, transistors? Only a small percentage of humanity are/were capable of doing any…

> Only a small percentage of humanity are/were capable of doing any of these. And they tend to be the best of the best in their respective fields.

Sure, agreed, but the difference between a small percentage and zero percentage is infinite.

Re: A new Google model is nearly perfect on automated handwriting recognition

#206

My task today for LLMs was "can you tell if this MRI brain scan is facing the normal way", and the answer was: no, absolutely not. Opus 4.1 succeeds more than chance, but still not nearly often enough to be useful. They all cheerfully hallucinate the wrong answer, confidently explaining the anatomy they are looking for, but wrong. Maybe Gemini 3 will pull it off. Now, Claude did vibe code a fairly accurate solution t…

That's fairly unfair comparison. Did you include in the prompt a basic set of instructions about which way is "correct" and what to look for?

Re: A new Google model is nearly perfect on automated handwriting recognition

#207

> In tabulating the “errors” I saw the most astounding result I have ever seen from an LLM, one that made the hair stand up on the back of my neck. Reading through the text, I saw that Gemini had transcribed a line as “To 1 loff Sugar 14 lb 5 oz @ 1/4 0 19 1”. If you look at the actual document, you’ll see that what is actually written on that line is the following: “To 1 loff Sugar 145 @ 1/4 0 19 1”. For those unawa…

I implemented a receipt scanner to Google Sheet using Gemini Flash.

The fact that it is ”intelligent" it's fine for some things.

For example I created structured output schema that had a field "currency" with the 3 letter format (USD, EUR...). So I scanned a receipt from some shop in Jakarta and it filled that field with IDR (Indonesian Rupiah). It inferred that data because of the city name on the receipt.

Would it be better for my use case that it would have returned no data for the currency field? Don't think so.

Note: if needed maybe I could have changed the prompt to not infer the currency when not explicitly listed on the receipt.

Re: A new Google model is nearly perfect on automated handwriting recognition

#208

I really hope they have because I’ve also been experimenting with LLMs to automate searching through old archival handwritten documents. I’m interested in the Conquistadors and their extensive accounts of their expeditions, but holy cow reading 16th century handwritten Spanish and translating it at the same time is a nightmare, requiring a ton of expertise and inside field knowledge. It doesn’t help that they were of…

I'm skeptical because my entire identity is basically built around being a software engineer and thinking my IQ and intelligence is higher than other people. If this AI stuff is real then it basically destroys my entire identity so I choose the most convenient conclusion. Basically we all know that AI is just a stochastic parrot autocomplete. That's all it is. Anyone who doesn't agree with me is of lesser intelligenc…

This kind of comment certainly shows that no organic stochastic parrots post to hn threads!

Re: A new Google model is nearly perfect on automated handwriting recognition

#209

Earlier quoted context omitted.

I remain confused but still somewhat interested as to a definition of "novel", given how often this idea is wielded in the AI context. How is everyone so good at identifying "novel"? For example, I can't wrap my head around how a) a human could come up with a piece of writing that inarguably reads "novel" writing, while b) an AI could be guaranteed to not be able to do the same, under the same standard.

Generally novel either refers to something that is new, or a certain type of literature. If the AI is generating something functionally equivalent to a program in its training set (in this case, dozens or even hundreds of such programs) then it by definition cannot be novel.

OK, but by that definition, how many human software developers ever develop something "novel"? Of course, the "functionally equivalent" term is doing a lot of heavy lifting here: How equivalent? How many differences are required to qualify as different? How many similarities are required to qualify as similar? Which one overrules the other? If I write an app that's identical to Excel in every single aspect except that instead of a Microsoft Flight Simulator easter egg, there's a different, unique, fully playable game that can't be summed up with any combination of genre lables, is that 'novel'?

Re: A new Google model is nearly perfect on automated handwriting recognition

#210

> In tabulating the “errors” I saw the most astounding result I have ever seen from an LLM, one that made the hair stand up on the back of my neck. Reading through the text, I saw that Gemini had transcribed a line as “To 1 loff Sugar 14 lb 5 oz @ 1/4 0 19 1”. If you look at the actual document, you’ll see that what is actually written on that line is the following: “To 1 loff Sugar 145 @ 1/4 0 19 1”. For those unawa…

[flagged]
Post reply on HN