Live data from Hacker News

AI’s Language Problem

technologyreview.com

151–160 of 244 posts

Re: AI’s Language Problem

#151

Earlier quoted context omitted.

I believe we can get AI from just text. Obviously that won't work for babies, because babies get bored quickly looking at text. AIs can be forced to read billions of words in mere hours! Look at word2vec. By using simple dimensionality reduction on the frequency that words that occur near to each other in news articles, it can learn really interesting things about the meaning of words. The famous example is the vecto…

> All of the information of our world is contained in text. This statement is false. There is a well known thought experiment called Mary’s Room the gist of which is that knowing all conceivable scientific knowledge about how humans perceive color is still not a substitute for being a human and perceiving the color red: https://philosophynow.org/issues/99/What_Did_Mary_Know The experience of seeing red is an example…

Mary's Room and qualia is totally irrelevant. I'm not asking if the computer will "feel" "redness", simply if it can pretend to do so through text. If it can talk about the color red, in a way indistinguishable from any other human talking about red.

In any case, at some level everything is symbols. A video is just a bunch of 1's and 0's, as is text, and everything else. A being raised on only text input would have qualia just like a being raised on video input. It would just be different qualia.

Re: AI’s Language Problem

#152

Earlier quoted context omitted.

The Turing Test equivalent in the game of Go is not beating the computer. Instead, it is: given a history of moves, can you determine whether or not a computer was playing? Determining whether or not AlphaGo was playing seems like a relatively easy task to me for the designers. Since they have access to the AlphaGo system, they can just calculate the probability that each move corresponds to one AlphaGo would make.

> Instead, it is: given a history of moves, can you determine whether or not a computer was playing? No, it's really not. For the Turing Test, the AI is meant to be adversarial—it's objective is to convince you that it is human. AlphaGo's objective isn't to "play like a human," it is to win. If they gave it an objective of playing like a human, I'm sure AlphaGo could play in a way that would be indistinguishable from…

> If they gave it an objective of playing like a human, I'm sure AlphaGo could play in a way that would be indistinguishable from a human.

It could just play unbelievably bad and appear like a beginner. That wouldn't prove intelligent.

> Peeking at the system/data is cheating

Someone ignorant of computers would hardly ever assume a machine. Of course the omission of this rule would leave someone smarter than the computer.

If you talk statistics, IE the machine has to convince only a fair share of humans, the definition of the threshold is a problem. Intelligence would depend on the development of the society. I thought this is about an intrinsic value.

It's an interesting thought experiment, but hardly conclusive, just observational.

Re: AI’s Language Problem

#153

Earlier quoted context omitted.

I believe we can get AI from just text. Obviously that won't work for babies, because babies get bored quickly looking at text. AIs can be forced to read billions of words in mere hours! Look at word2vec. By using simple dimensionality reduction on the frequency that words that occur near to each other in news articles, it can learn really interesting things about the meaning of words. The famous example is the vecto…

> All of the information of our world is contained in text. This statement is false. There is a well known thought experiment called Mary’s Room the gist of which is that knowing all conceivable scientific knowledge about how humans perceive color is still not a substitute for being a human and perceiving the color red: https://philosophynow.org/issues/99/What_Did_Mary_Know The experience of seeing red is an example…

The Mary's Room thought experiment is garbage, if you ask me. You can't just assume your hypothesis and then call the result truth.

If you assert that a person can understand everything there is to know about the color red and then still not understand what it is like to see red, you have either contradicted yourself or assumed dualism.

Re: AI’s Language Problem

#154

Earlier quoted context omitted.

I believe we can get AI from just text. Obviously that won't work for babies, because babies get bored quickly looking at text. AIs can be forced to read billions of words in mere hours! Look at word2vec. By using simple dimensionality reduction on the frequency that words that occur near to each other in news articles, it can learn really interesting things about the meaning of words. The famous example is the vecto…

I think if AGI were possible from the basic statistical NLP techniques outlined in most advanced NLP textbooks, it would have already happened a decade ago.

I'm not saying it is possible from just basic statistical NLP techniques. It may take much more advanced techniques. And it may take much more computing power than we have even now.

But I do believe it is possible, someday. Probably within our lifetime.

Re: AI’s Language Problem

#155
post #142

Earlier quoted context omitted.

It's probably a mistake to assume that just because it's the way we do it that it has to be the way machines do it. Although that's usually the initial assumption. In the early days of flight most attempts were based on birds, similarly submersible vehicles were based on fish. We know now it's better to use propellers. It could be we just haven't found what is analogous to a propeller for the AI world.

The only general intelligence we know of is us. It stands to reason that the first step towards creating AGI is to copy the one machine we know is capable of that type of processing. Why doesn't our research focus on understanding and copying biological brains? Numenta did, with good results, but it isn't an industry trend.

The steps to learning to make aircraft didn't come from understanding how birds flap their wings. I can't imagine the kind of intelligence we consider general will be from study of how humans biologically think.

Re: AI’s Language Problem

#156
post #113

Earlier quoted context omitted.

Computers are programmed only using text, even if the text has just two symbols. Sensor interfaces use digital signals, again symbol streams. That would beget the question if there could be machines mightier than a Turing Complete one. I'm sure that's missing your point.

If you could simply program a human's behavior into a machine, that would be fine, but humans aren't capable of encoding their neural circuitry in a programming language; in fact, the combined efforts of all humans have only begun to shed light on what human behavior is. As such, generating formalized information (that describes behavior) requires some process (other than human hands) to do it -- and in the case of a…

By Shannon's information theory, everything is bits of information. Completely irrelevant for the mechanism of a Turing Machine is, who writes the initial bits on the tape. For all I care, the world is the tape and the computer is the head being moved through it. It's the old fallacy seeing the brain as a computer, since we've build computers after structures from our brains. Hence I wondered, what more there should be.

Re: AI’s Language Problem

#157

Earlier quoted context omitted.

It can be the case if Chomsky was right, and Universal Grammar and other similar structures are a thing. That would mean that part of our ability to understand language comes from the particular structure of our brain (which everyone seems to by and large share). That would mean that some of our ability to understand language is genetic in nature, by whatever means genes direct the structure of brain development.

So if language comes from the structure of the brain, what would stop us from simulating that structure to give a machine mastery of language? And specifically what would imply that a machine which had some of that structure would need to learn by interaction as the top level comment suggests?

Nothing would stop us from simulating human brain-like (or analogously powerful) structures to build a machine that genuinely understands natural language. I'm arguing that we can't just learn those structures by statistical optimization techniques though.

If it turns out that the easiest, or even only means of doing this is by emulating the human brain, then it is entirely possible that we inherit a whole new set of constraints and dependencies such that world-simulation and an emobdied mind are required to make such a system learn. If this turns out not to be the case, that there's some underlying principle of language we can emulate (the classic "airplanes don't fly like birds" argument) then it may be the case that text is enough. But that's in the presence of a new assumption, that our system came pre-equipped to learn language, and didn't manufacture an understanding from whole cloth. That the model weights were pre-initialized to specific values.

Re: AI’s Language Problem

#158

Earlier quoted context omitted.

> What is the meaning of a word, if not the relationship it has to other words? There's also the relationship it has with the world.

Well in my example the AI doesn't have to interact with the world at all. To pass the Turing test simply requires imitating a human, predicting what words they would say. You only need to know the relationships between words.

"Is it day or night?"

If literally the only thing you know is the relationship between words, but you have a perfect knowledge of the relationship between words, you'll quickly determine that "Day" and "Night" are both acceptable answers, and have no means of determining which is the right one. At the very minimum, you need a clock, and an understanding of the temporal nature of your training set, to get the right one.

Re: AI’s Language Problem

#159

Earlier quoted context omitted.

It's learning the meaning of words, and the relationships between them. Word2vec is definitely an impressive algorithm. But at the end of the day, it's just a tool that cranks out a fine-grained clustering of words based on (a proxy measure for) contextual similarity (or rather: an embedding in a high-dimensional space, which implicitly allows the words to be more easily clustered). And yes, some additive relations b…

Word2vec may be crude, but it demonstrates that you can learn non-trivial relationships between words with even such a simple algorithm. What is the meaning of a word, if not the relationship it has to other words? Gender was just an example. There are lots of semantic information learned by word2vec, and the vectors have shown to be useful in text classification and other uses. It can learn subtle stuff, like the re…

But it and methods like it are still very limited in what they can learn. For example, they can't learn relations involving antonyms. They can't tell apart hot from cold or big from small.

Re: AI’s Language Problem

#160

Earlier quoted context omitted.

It can be the case if Chomsky was right, and Universal Grammar and other similar structures are a thing. That would mean that part of our ability to understand language comes from the particular structure of our brain (which everyone seems to by and large share). That would mean that some of our ability to understand language is genetic in nature, by whatever means genes direct the structure of brain development.

But I don't see any reason a "universal grammar" couldn't be learned. It may take something more complicated than ANNs, of course. But it would be really weird if there was a pattern in language that was so obfuscated it couldn't be detected at all.

it comes down to the limits of available Information with a capital 'I'. If you're working within the encoding system (as you're recommending here with the "all the text in the world" approach), then in order to learn the function that's generating this information, the messages that you're examining have a minimum amount of information they can convey. There needs to be enough visible structure purely within the context of the messages themselves to make the underlying signal clear.

I don't think it's so weird to imagine that natural language really doesn't convey a ton of explicit information on its own. Sure, there's some there, enough that our current AI attempts can solve little corners of the bigger problem. But is it so strange to imagine that the machinery of the human brain takes lossy, low-information language and expands, extrapolates, and interprets it so heavily so as to make it orders of magnitude more complex than the lossy, narrow channel through which it was conveyed? That the only reason we're capable of learning language and understanding eachother (the times we _do_ understand eachother) is because we all come pre-equipped with the same decryption hardware?

Post reply on HN