Live data from Hacker News

AI’s Language Problem

technologyreview.com

211–220 of 244 posts

Re: AI’s Language Problem

#211

Earlier quoted context omitted.

Helen Keller's life argues against the proposition that a machine requires the same sort of interaction with its environment that the average human experiences, before it can achieve intelligence.

She was blind and deaf, but she still had enormous amounts of tactile information and the ability to physically interact with her environment. Moreover, those sensory elements were critical for her to finally be able to start learning language--she finally caught on that the signs another person was making in one hand represented the water being run over her other hand. And she didn't make that breakthrough until she…

So Keller's "enormous amounts of tactile information and the ability to physically interact with her environment" means that her achievements were not that remarkable? Speaking personally, if I were to lose my sight and hearing, I imagine I would find life extremely daunting, even though I have what seems to me the very considerable advantage of having learned language.

Anne Sullivan's equally remarkable (IMHO) role as a teacher is not really an issue here, as training is also an option for AI, though it might be evidence in a rather different discussion about whether unsupervised learning alone, particularly as practiced today, is likely to get us to AI (clearly, the evolution of intelligence can be cast as unsupervised learning, but that is a very long and uncertain process...)

Re: AI’s Language Problem

#212
post #206

Earlier quoted context omitted.

This is why Turing invented his famous test. At the time people were arguing about what it would take to prove a machine is intelligent. People argued that the internal properties of the machine mattered, that it needed to feel emotions and qualia just like humans. But Turing argued that if the machine could be shown to do the same tasks that humans can do, and act indistinguishable from a real human, then surely it'…

Exactly. The computer will fail such a test where such tests can only be passed if and only if the agent under test experiences subjective states. Since humans can always craft this class of tests, and the computer cannot pass it, it will always fail the "Turing Test".

What test could possibly test for subjective states? You can ask the computer how it feels, and it can just lie, or predict what a human would say if asked the same question. There's no way to know what the computer actually feels, and it doesn't really matter for this purpose.

Re: AI’s Language Problem

#213
post #206

Earlier quoted context omitted.

Exactly. The computer will fail such a test where such tests can only be passed if and only if the agent under test experiences subjective states. Since humans can always craft this class of tests, and the computer cannot pass it, it will always fail the "Turing Test".

What test could possibly test for subjective states? You can ask the computer how it feels, and it can just lie, or predict what a human would say if asked the same question. There's no way to know what the computer actually feels, and it doesn't really matter for this purpose.

The easy answer is this: these tests exist. Since no computer put to the turing test has passed, simply look up the test and observe how humans have induced the computer to fail.

In practice, a good class of tests to use is a test that must evoke an emotional response to produce a sensical answer. An example is art interpretation. Questions involving allegory. Interpret a poem etc.

Important to note that whatever the challenge is, it must always be a new example - as in never been seen before. Anything that is already in the existing corpus, the computer can simply look up what is already out there. In other words, there is no one concrete thing you can use again and again repeatedly.

Example of test that would foil a computer: A personally written poem and having discussion about it.

Re: AI’s Language Problem

#214

Earlier quoted context omitted.

Learning the statistics of language is not going to tell the ML model anything about the underlying stuff to which the language actually refers . It will need actual "sense-data" to do that. For instance, to get a model that generates image captions, you need to train it with actual images. If the much-vaunted "general intelligence" consists in both vague and precise causal reasoning and optimal control with respect…

There is nothing magical about "sense data". A video is just a bunch of 1's and 0's, just like text data. A model of video data is not superior in any way to one of text data, they are just different. The internet is so large and so comprehensive (especially if you include digitized books and papers, e.g. libgen or google books) that I doubt any important information that can be learned through video data, can't be o…

>A video is just a bunch of 1's and 0's, just like text data. A model of video data is not superior in any way to one of text data, they are just different.

Uhhhh... there's this thing called entropy. A corpus of video data is vastly higher in entropy content than text data, and probably a good deal more tightly correlated too, making it much easier to learn from.

Remember, for a human, the whole point of speech and text is to function as a robust, efficient code (in the information-theoretic sense) for activating generative models we already have in our heads when we acquire the words. This is why we have such a hard time with very abstract concepts like mathematical ones: the causal-role concepts (easy to theorize how to encode those as generative models) are difficult to acquire from nothing but the usage statistics of the words and symbols, in contrast to "concrete", sense-grounded concepts, which have large amounts of high-dimensional data to fuel the Blessing of Abstraction.

Nevermind, I should probably just get someone to let me into a PhD program so I can publish this stuff. If only they'd consider the first paper novel enough already!

Re: AI’s Language Problem

#215

Earlier quoted context omitted.

There is nothing magical about "sense data". A video is just a bunch of 1's and 0's, just like text data. A model of video data is not superior in any way to one of text data, they are just different. The internet is so large and so comprehensive (especially if you include digitized books and papers, e.g. libgen or google books) that I doubt any important information that can be learned through video data, can't be o…

>A video is just a bunch of 1's and 0's, just like text data. A model of video data is not superior in any way to one of text data, they are just different. Uhhhh... there's this thing called entropy . A corpus of video data is vastly higher in entropy content than text data, and probably a good deal more tightly correlated too, making it much easier to learn from. Remember, for a human, the whole point of speech and…

I think you mean redundancy. And yes videos are highly redundant. But I don't see how that's any kind of advantage. Text has all the relevant information contained within it, with a ton of irrelevant information discarded. But there are still plenty of learnable patterns. Even trivial algorithms like word2vec can glean a huge amount of semantic information (much easier than is possible with video, currently.)

I don't know if humans have generative models in their heads. There are people who have no ability to form mental images, and they function fine. Regardless, an AI should be able to get around that by learning our common patterns. It need not mimic our internal states, only our external behavior.

Re: AI’s Language Problem

#217
post #72

Earlier quoted context omitted.

Sorry, I've edited my original comment to be clearer. What I really meant is that there is wide tolerance of noise in those domains. "How long does stars last" has a completely different meaning than "How long do stars last" - not tolerant of noise.

If an 6th grader asks their science teacher "How long does stars last?" / "How long stars last?" /"How long do stars last?" / "How old do stars get?" / "Stars, how old can they get?" / ... In similar context they probably end up parsed to the same question assuming correct inflection, posture, etc. Spoken conversations are messy, but they also have redundancy and pseudo checksum's. Written language tends to be more f…

All those sentences sound (mentally) a lot differently. Some of those sentences give you the impression the speaker is an idiot, for example.

Re: AI’s Language Problem

#218
post #8

Deep learning has succeeded tremendously with perception in domains that tolerate lots of noise (audio/visual). Will those successes continue with perception in domains that are not noisy (language) and inference/control , which the article touches on? I think it really is unclear whether those challenges will require fundamental developments or just more years of incremental improvement. If fundamental developments…

There is evidence that language is fairly smooth though. For example, we can extract e.g. the gender vector from a word embedding space that is learned by a recurrent neural network. That seems to hint at the possibility that words, sentences and concepts live in smooth, high-dimensional manifold that makes them learnable for us in the first place (because in that case they can be learned by small local improvements…

The very method of using a word embedding space assumes the manifold is smooth, so the fact that vectors extracted from a method that assumes a smooth manifold, are in fact on a smooth manifold, is just circular and not evidence of anything.

Re: AI’s Language Problem

#219

Earlier quoted context omitted.

>A video is just a bunch of 1's and 0's, just like text data. A model of video data is not superior in any way to one of text data, they are just different. Uhhhh... there's this thing called entropy . A corpus of video data is vastly higher in entropy content than text data, and probably a good deal more tightly correlated too, making it much easier to learn from. Remember, for a human, the whole point of speech and…

I think you mean redundancy. And yes videos are highly redundant. But I don't see how that's any kind of advantage. Text has all the relevant information contained within it, with a ton of irrelevant information discarded. But there are still plenty of learnable patterns. Even trivial algorithms like word2vec can glean a huge amount of semantic information (much easier than is possible with video, currently.) I don't…

>I think you mean redundancy.

No, I meant a pair of specific information-theoretical quantities I've been studying.

>And yes videos are highly redundant. But I don't see how that's any kind of advantage.

Representations are easier to learn for highly-correlated data. Paper forthcoming, but conceptually so obvious that quantifying it is (apparently) non-novel.

>I don't know if humans have generative models in their heads.

The best available neuroscience and computational cognitive science says we do.

>Regardless, an AI should be able to get around that by learning our common patterns. It need not mimic our internal states, only our external behavior.

Our external behavior is determined by the internal states, insofar as those internal states are functions which map sensory (including proprioceptive and interoceptive) statistics to distributions over actions. If you want your robots to function in society, at least well enough to take it over and kill everyone, they need a good sense of context and good representations for structured information, behavior, and goals. Further, most of the context to our external behavior is nonverbal. You know how it's difficult to detect sarcasm over the internet? That's because you're trying to serialize a tightly correlated high-dimensional data-stream into a much lower-dimensional representation, and losing some of the variance (thus, some of the information content) along the way. Humans know enough about natural speech that we can usually, mostly reconstruct the intended meaning, but even then, we've had to develop a separate art of good writing to make our written representations conducive to easy reconstruction of actual speech.

Deep learning can't do this stuff right now. Objectively speaking, it's kinda primitive, actually. OTOH, realizing the implications of our best theories about neuroscience and information theory as good computational theories for cognitive science and ML/AI is going to take a while!

Re: AI’s Language Problem

#220

Earlier quoted context omitted.

> All of the information of our world is contained in text Communicate the concept of "green" to me in text. The sound of a dog barking, a motor turning over, a sonic boom, or the experience of a Doppler shift. Beethoven's symphony. Sour. Sweet. What does "mint" taste like? Shame. Merit. Learn facial and object recognition via text. Vertigo. Tell a boxer how to box by reading? Hand eye coordination, bodies in 3 dimen…

Are blind or deaf people not intelligent? But if you must pretend to be sighted and hearing, there are many descriptions of green, of dogs barking, of motors, etc, scattered through the many books written in English (and other languages.) Are these descriptions perfect? Maybe not. But they are sufficient to mimic or communicate with humans through text. It's sufficient to beat a Turing test, to answer questions intel…

Yes they are. However, is a blind, deaf, person with absolutely no motor control, no sense of touch, and no proprioception intelligent? Unclear. They certainly have no language faculties.
Post reply on HN