Live data from Hacker News

A new Google model is nearly perfect on automated handwriting recognition

generativehistory.substack.com

271–280 of 328 posts

Re: A new Google model is nearly perfect on automated handwriting recognition

#271
post #58

I've been complaining on hn for some time now that my only real test of an LLM is that it can help my poor wife with her research, she spends all day every day in small town archives pouring over 18th century American historical documents. I thought maybe that day had come, I showed her the article and she said "good for him I'm still not transcribing important historical documents with a chat bot and nor should he"…

> ...of an LLM is that it can help my poor wife with her research, she spends all day every day in small town archives pouring over 18th century American historical documents.

> I'm still not transcribing important historical documents with a chat bot and nor should he

Doesn't sound like she's interested in technology, or wants help.

Re: A new Google model is nearly perfect on automated handwriting recognition

#272

Earlier quoted context omitted.

> This is the magic argument reskinned. Transformers aren’t copying strings, they’re constructing latent representations that capture relationships, abstractions, and causal structure because doing so reduces loss. Sure (to the second part), but the latent representations aren't the same as a humans. The human's world that they have experience with, and therefore representations of, is the real word. The LLM's world…

>So, yes, there is plenty of depth and nuance to the internal representations of an LLM, but no logical reason to think that the "world model" of an LLM is similar to the "world model" of a human since they live in different worlds, and any "understanding" the LLM itself can be considered as having is going to be based on it's own world model. So LLMs and Humans are different and have different sensory inputs. So wha…

Regarding cats on mats ...

If you ask a human to complete the phrase "the cat sat on the", they will probably answer "mat". This is memorization, not understanding. The LLM can do this too.

If you just input "the cat sat on the" to an LLM, it will also likely just answer "mat" since this is what LLMs do - they are next-word input continuers.

If you said "the sat sat on the" to a human, they would probably respond "huh?" or "who the hell knows!", since the human understands that cats are fickle creatures and that partial sentences are not the conversational norm.

If you ask an LLM to explain it's understanding of cats, it will happily reply, but the output will not be it's own understanding of cats - it will be parroting some human opinion(s) it got from the training set. It has no first hand understanding, only 2nd hand heresay.

Re: A new Google model is nearly perfect on automated handwriting recognition

#273

Earlier quoted context omitted.

This is knock against you at all , but in a naive attempt to spare someone else some time: remember that based on this definition it is impossible for an LLM to do novel things and more importantly , you're not going to change how this person defines a concept as integral to one's being as novelty. I personally think this is a bit tautological of a definition, but if you hold it, then yes LLMs are not capable of anyt…

I think you should reverse the question, why would we expect LLMs to even have the ability to do novel things? It is like expecting a DJ remixing tracks to output original music. Confusing that the DJ is not actually playing the instruments on the recorded music so they can't do something new beyond the interpolation. I love DJ sets but it wouldn't be fair to the DJ to expect them to know how to play the sitar becaus…

kid koala does jazz solos on a disk of 12 notes, jumping the track back and forth to get different notes.

i think that, along with the sitar player are still interpolating. the notes are all there on the instrument. even without an instrument, its still interpolating. the space that music and aound can be in is all well known wave math. if you draw a fourier transform view, you could see one chart with all 0, and a second with all +infinite, and all music and sound is gonna sit somewhere between the two.

i dont know that "just interpolation" is all that meaningful to whether something is novel or interesting.

Re: A new Google model is nearly perfect on automated handwriting recognition

#274

Earlier quoted context omitted.

This is knock against you at all , but in a naive attempt to spare someone else some time: remember that based on this definition it is impossible for an LLM to do novel things and more importantly , you're not going to change how this person defines a concept as integral to one's being as novelty. I personally think this is a bit tautological of a definition, but if you hold it, then yes LLMs are not capable of anyt…

That is not strictly true, because being able to transfer the style of Van Gogh onto an arbitrary photographic scene is novel in a sense, but it is interpolative. Mashups are not purely derivative: the choice of what to mash up carries novelty: two (or more) representations are mashed together which hitherto have not been. We cannot deny that something is new.

I'm saying their comment is calling that not something new.

I don't agree, but by their estimation adding things together is still just using existing things.

Re: A new Google model is nearly perfect on automated handwriting recognition

#276

Earlier quoted context omitted.

Cept its not made Linux (in the absence of it). At any point prior to the final output it can garner huge starting point bias from ingested reference material. This can be up to and including whole solutions to the original prompt minus some derivations. This is effectively akin to cheating for humans as we cant bring notes to the exam. Since we do not have a complete picture of where every part of the output comes f…

> Cept its not made Linux (in the absence of it). Neither did you (or I). Did you create anything that you are certain your peers would recognize as more "novel" than anything a LLM could produce?

>Neither did you (or I).

Not that specifically but I certainly have the capability to create my own OS without having to refer to the source code of existing operating systems. Literally "creating a linux" is a bit on the impossible side because it implies compatibility with an existing kernel despite the constraints prohibiting me from referring to the source of that existing kernel (maybe possible if i had some clean-room RE team that would read through the source and create a list of requirements without including any source).

If we're all on the same page regarding the origins of human intelligence (ie, that it does not begin with satan tricking adam and eve into eating the fruit of a tree they were specifically instructed not to touch) then it necessarily follows that any idea or concept was new at some point and had to be developed by somebody who didn't already have an entire library of books explaining the solution at his disposal.

For the Linux thought-experiment you could maybe argue that Linux isn't totally novel since its creator was intentionally mimicking behavior of an existing well-known operating system (also iirc he had access to the minix source) and maybe you could even argue that those predecessors stood on the shoulders of their own proverbial giants, but if we keep kicking the ball down the road eventually we reach a point where somebody had an idea which was not in any way inspired by somebody else's existing idea.

The argument I want to make is not that humans never create derivative or unoriginal works (that obviously cannot be true) but that humans have the capability to create new things. I'm not convinced that LLMs have that same capability; maybe I'm wrong but I'm still waiting to see evidence of them discovering something new. As I said in another post, this could easily be demonstrated with a controlled experiment in which the model is bootstrapped with a basic yet intentionally-limited "education" and then tasked with discovering something already known to the experimenters which was not in its training set.

>Did you create anything that you are certain your peers would recognize as more "novel" than anything a LLM could produce?

Yes, I have definitely created things without first reading every book in the library and memorizing thousands of existing functionally-equivalent solutions to the same problem. So have you so long as I'm not actually debating an LLM right now.

Re: A new Google model is nearly perfect on automated handwriting recognition

#277

Earlier quoted context omitted.

>So, yes, there is plenty of depth and nuance to the internal representations of an LLM, but no logical reason to think that the "world model" of an LLM is similar to the "world model" of a human since they live in different worlds, and any "understanding" the LLM itself can be considered as having is going to be based on it's own world model. So LLMs and Humans are different and have different sensory inputs. So wha…

Regarding cats on mats ... If you ask a human to complete the phrase "the cat sat on the", they will probably answer "mat". This is memorization, not understanding. The LLM can do this too. If you just input "the cat sat on the" to an LLM, it will also likely just answer "mat" since this is what LLMs do - they are next-word input continuers. If you said "the sat sat on the" to a human, they would probably respond "hu…

>If you said "the sat sat on the" to a human, they would probably respond "huh?" or "who the hell knows!", since the human understands that cats are fickle creatures and that partial sentences are not the conversational norm.

I'm not sure what you're getting at here ? You think LLMs don't similarly answer 'What are you trying to say?'. Sometimes I wonder if the people who propose these gotcha questions ever bother to actually test them on said LLMs.

>If you ask an LLM to explain it's understanding of cats, it will happily reply, but the output will not be it's own understanding of cats - it will be parroting some human opinion(s) it got from the training set. It has no first hand understanding, only 2nd hand heresay.

Again, you're not making the distinction you think you are. Understanding from '2nd hand heresay' is still understanding. The vast majority of what humans learn in school is such.

Re: A new Google model is nearly perfect on automated handwriting recognition

#278

Earlier quoted context omitted.

You are right to be skeptical. There are plenty of so called windows(or other) web 'os' clones. There were a couple of these posted on HN actually this very year. Here is one example I google dthat was also on HN : https://news.ycombinator.com/item?id=44088777 This is not an OS as in emulating a kernel in javascript or wasm, this is making a web app that looks like the desktop of an OS. I have seen plenty such projec…

Its always amusing when "an app like windows xp" considered hard or challenging somehow. Literally the most basic html/css, not sure why it is even included in benchmarks.

While it is obviously much easier than creating a real OS, some people have created desktop managers web apps, with resizeable and movable windows, apps such as terminals, nodepads, file explorer etc.

This is still a challenging task and requires lots of work to get this far.

Re: A new Google model is nearly perfect on automated handwriting recognition

#279

Earlier quoted context omitted.

>So, yes, there is plenty of depth and nuance to the internal representations of an LLM, but no logical reason to think that the "world model" of an LLM is similar to the "world model" of a human since they live in different worlds, and any "understanding" the LLM itself can be considered as having is going to be based on it's own world model. So LLMs and Humans are different and have different sensory inputs. So wha…

> You think dolphins and orcas are not intelligent and don't understand things ? Not sure where you got that non-secitur from ... I would expect most animal intelligence (incl. humans) to be very similar, since their brains are very similar. Orcas are animals. LLMs are not animals.

Orca and human brains are similar, in the sense we have a common ancestor if you look back far enough, but they are still very different and focus on entirely different slices of reality and input than humans will ever do. It's not something you can brush off if you really believe in input supremacy so much.

From the orca's perspective, many of the things we say we understand are similarly '2nd hand hearsay'.

Post reply on HN