Live data from Hacker News

Machine learning won't solve natural language understanding

thegradient.pub

11–20 of 193 posts

Re: Machine learning won't solve natural language understanding

#11
I didn't get beyond his argument that ML is compression while NLU is decompression but that doesn't seem to be right. ML is often used to decompress data such as increasing image resolution. Of course it needs more data than just the compressed form, but only for training. For inference, of course you can use ML to add assumed common knowledge information.

Re: Machine learning won't solve natural language understanding

#12
Going through this, it has the old "A full understanding of an utterance or a question requires understanding the one and only one thought that a speaker is trying to convey. " claim, which continues to not make any sense, because obviously people don't do that; as much as I would like to be understood in precisely the way I mean, down to the most subtle nuance/shade of meaning and connotation, at least much of the time, this is not something we can actually get across in an at all reasonable amount of time.

Also, claiming that natural language is infinite, if taken literally, would imply a large and contrary to the common consensus claim about physics, contradicting the Bekenstein bound and all that.

But one thing which seemed, at least initially, like a point that could have some merit, was the point about compression vs decompression.

But the alleged syllogism about it, is, pretending to be much more formal/rigorous than it is, and is also kind of nonsense? Or, like, it conflates "NLU is about decompression" with, "NLU \equiv not COMP" which I assume is meant to mean, -- well, actually, I'm not sure what it is supposed to mean. Initially I thought it was supposed to mean "NLU is nonequivalent to compression", but if so, it should be written as like, "NLU \not\equiv COMP" (where \not\equiv is the struckthrough version of the \equiv symbol) , but if it is supposed to mean "NLU is equivalent to the inverse or opposite of compression" (which I suppose better fits the text description on the right better), then I don't think "not" is the appropriate way to express that. And, if by "not" the author really means "the inverse of", then, well, there's nothing wrong with something being equivalent to its own inverse! Nor, does something being equivalent to the inverse of something else imply that it is "incompatible" with it.

For something talking about communicating ideas and these ideas being understood precisely by the recipient of the message, the author sure did not work to communicate precisely.

The value in formalization comes not in its trappings, but in actually being careful and precise, etc., not merely pretending to be.

The part on intensional equality vs extensional equality was interesting, but the claim that neural networks cannot represent intension is, afaict, not given any justification (other than just "because they are numeric").

Re: Machine learning won't solve natural language understanding

#13

> In other words, we must get, from a multitude of possible interpretations of the above question, the one and only one meaning that, according to our commonsense knowledge of the world, is the one thought behind the question some speaker intended to ask. But, can humans do this? I think not; I still disagree with the author about what "Do we have a retired BBC reporter that was based in an East European country duri…

Well, that's true, but I don't think we can pretend that computers are anywhere near as good at humans at extracting the intended meaning from human utterances.

Re: Machine learning won't solve natural language understanding

#14
post #10

> In other words, we must get, from a multitude of possible interpretations of the above question, the one and only one meaning that, according to our commonsense knowledge of the world, is the one thought behind the question some speaker intended to ask. But, can humans do this? I think not; I still disagree with the author about what "Do we have a retired BBC reporter that was based in an East European country duri…

> Even humans are only probably approximately correct. This is very true, more true than we realize. Notice how much more "could you repeat that?" we have with masks on. It's not JUST the mild muffling of the speaker's voice, it's not seeing their lips move. We're all lip readers to a small degree, and it helps inform our decoding to see the lips. Fff and th sound similar but look very different. Even without that, t…

That's a different problem, isn't it? That's more about transcription -- getting the speech into words -- than about what he's talking about, making sense of the words once you have them.

Re: Machine learning won't solve natural language understanding

#15
Hunter-gatherer's brain did not evolve to facilitate unambiguous thought transmission.

Speech evolved to be very visual, spacial, quantitative and political. Ability to lie efficiently is evolutionary trait. Sometimes, we don't even need words for that. And sometimes we lie by pronouncing exclusively true words and sentences. Ambiguity of speech always was a feature.

None of that makes work of NL researcher easier, of course.

Understanding a sentence is not like decoding or decompressing, it is more like trying to guess what is utterer up to, politically, and if he is a friend. And only then there is deciding where to steer according to what he says. And for that we sometimes should start decoding message, but only with sender's goals firmly materialized in mind.

Re: Machine learning won't solve natural language understanding

#16

I didn't get beyond his argument that ML is compression while NLU is decompression but that doesn't seem to be right. ML is often used to decompress data such as increasing image resolution. Of course it needs more data than just the compressed form, but only for training. For inference, of course you can use ML to add assumed common knowledge information.

> For inference, of course you can use ML to add assumed common knowledge information.

I think this is easier said than done, and hasn't yet been accomplished in any general sense.

Re: Machine learning won't solve natural language understanding

#17
post #2

An armchair effort to redefine the goalposts and judge NLP, but proof is in the pudding. For now, large language models are the best flavor. NLP models are already useful even in this early stage.

> An armchair effort Given the authors credentials and publication history [1] it's a bit disingenuous to call this an 'armchair effort'. [1] https://scholar.google.com/citations?user=i5sEc1YAAAAJ&hl=en...

My h index is twice as high and that's a terrible h index still lol

Re: Machine learning won't solve natural language understanding

#18

> In other words, we must get, from a multitude of possible interpretations of the above question, the one and only one meaning that, according to our commonsense knowledge of the world, is the one thought behind the question some speaker intended to ask. But, can humans do this? I think not; I still disagree with the author about what "Do we have a retired BBC reporter that was based in an East European country duri…

> But, can humans do this? I think not;

> Even humans are only probably approximately correct.

Fair point. But how about this:

It's true that what John would do with the sentence is technically an "approximation" of what Alice would do because they have slightly different understandings of correct behavior. However, for humans to do what they do, they still do build an absolutely correct model of meaning in their mind wrt their (subjective) notion of correctness.

This may sound like an obtuse play with words but the point is that to even attempt to do the right kind of reasoning in NLU, you need a different framework than PAC. You can't for example approximate whether "during the Cold War" qualifies "was based in" or qualifies "an Eastern European country". You just need to decide. And once you decide, you have an absolute correct interpretation, not an approximate one.

EDIT: wording.

Re: Machine learning won't solve natural language understanding

#19
post #12

Going through this, it has the old "A full understanding of an utterance or a question requires understanding the one and only one thought that a speaker is trying to convey. " claim, which continues to not make any sense, because obviously people don't do that; as much as I would like to be understood in precisely the way I mean, down to the most subtle nuance/shade of meaning and connotation, at least much of the t…

> Also, claiming that natural language is infinite, if taken literally, would imply a large and contrary to the common consensus claim about physics, contradicting the Bekenstein bound and all that.

Natural language is infinite in the pretty straightforward sense that, say, chess is infinite (there is an infinite number of valid chess games - if you ignore arbitrary restrictions such as the 50 move rule). This of course doesn’t mean that a chess computer has to be infinitely large or violate any known laws of physics. Similarly, practical C compilers can exist despite their being an infinite number of valid C programs.

Re: Machine learning won't solve natural language understanding

#20
I think that's very true and it's maybe even more clear when you consider mathematics.

You can maybe imitate but not effectively learn mathematics empirically. There is an infinite number of mathematical expressions or sequences that can be generated, so learning can never be done, you cannot compress yourself to mathematical understanding. (which is obvious if you try to feed language models simple arithmetic, they can maybe do 5+5 because it shows up somewhere in the data, but then they can't do 3792 + 29382, hence they do not understand anything about addition at all).

The correct way to mathematical understanding is decompressing, understanding the fundamental axioms of mathematics and internal relationships of mathematical objects (comparable to the semantic meaning behind language artifacts), and then expanding them.

Post reply on HN