Live data from Hacker News

Large models of what? Mistaking engineering achievements for linguistic agency

arxiv.org

61–70 of 162 posts

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#61

The authors of this paper are just another instance of the AI hype being used by people who have no connection to it, to attract some kind of attention. "Here is what we think about this current hot topic; please read our stuff and cite generously ..." > Language completeness assumes that a distinct and complete thing such as `a natural language' exists, the essential characteristics of which can be effectively and c…

They are two researchers/assistant professors working with cognitive science, psychology, and trustworthy AI. The paper is peer reviewed and has been accepted for publication in the Journal of Language Sciences. You should publish your critique of their research in that same journal. P.s. if you find any grave mistakes, you can contact the editor in chief, who happens to be a linguist.

The "efficient journal hypothesis" -- if something is written in a paper in a journal, then it's impossible for anyone to know any better, since if they knew better, they would already have published the correction in a journal.

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#62
post #28

Earlier quoted context omitted.

I'm not so sure that view is very widespread amongst people familiar with how LLMs work. Certainly they become more capable with parameters and data, but there are fundamental things that can't be overcome with a basic model and I don't think anyone is seriously arguing otherwise. For instance LLMs are pretty much stateless without their context window. If you treat the raw generated output as the first and final res…

Has anyone done the thing, and achieved some interesting results? Seems a pretty obvious thing to try, but I never heard of anything like it.

I've seen automated AI agents that can spend time reflecting on themselves in a feedback loop. The model alters its state over time and can call APIs.

You could equate saying something "aloud" to calling an API.

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#63

Earlier quoted context omitted.

LLMs do contain conceptual representations and LLMs are capable of abstract reasoning. This is trivially provable by asking them to reason about something that is a) purely abstract and b) not in the training data, e.g. "All floots are gronks. Some gronks are klorps. Are any floots klorps?" Any of the leading LLMs will correctly answer questions of this type much more often than chance.

That is not an example of a LLM being capable of abstract reasoning. Changing the question from "What is the capital of United States?" which is easily answerable to something completely abstract and "not in the training model" doesn't change that LLM's are just very advanced text prediction, and always will be. The nature of their design means they are incapable of AGI.

> LLM's are just very advanced text prediction, and always will be

How do you predict the next word in answering an abstract logic question without being capable of abstract reasoning, though?

In some sense it probably is possible, but this is a gaping flaw in your argument. A sufficiently advanced text prediction process has to encompass the process of abstract reasoning. The text prediction problem is necessarily a superset of the abstract reasoning problem. Ie, in the limit text prediction is fundamentally harder than abstract reasoning.

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#64

How would the authors consider a paralyzed individual who can only move their eyes since birth? That person can learn the same concepts as other humans and communicate as richly (using only their eyes) as other humans. Clearly, the paper is viewing the problem very narrowly.

> ...a paralyzed individual who can only move their eyes since birth...

I don't think such an individual is possible.

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#65

Earlier quoted context omitted.

Huh, almost right. ("possible, but not guaranteed?" it's necessarily true. That whole sentence was a waste of space, and wrong.) Edit: I mean "if there is any overlap", it's necessarily true. I should have quoted the whole thing.

Nope, ChatGPT was right, the answer is indeterminable. The klorps that are gronks could be a wholly distinct subset to the klorps that are floots. It also correctly evaluates "All gronks are floots. Some gronks are klorps. Are any floots klorps?", to which the answer is definitively yes.

> The klorps that are gronks could be a wholly distinct subset to the klorps that are floots.

So? It's still the case that "if there is any overlap between floots and klorps," it is "guaranteed, that some floots are klorps." It's tautological.

Unless there's a way to read "overlap" so that it doesn't mean "some of one category are also in the other category, and vice versa"?

Oh, when I said "it's necessarily true" I was refering to this last sentence of the output, not the question posed in the input. Hence we are at cross purposes I think.

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#66
post #28

Earlier quoted context omitted.

I'm not so sure that view is very widespread amongst people familiar with how LLMs work. Certainly they become more capable with parameters and data, but there are fundamental things that can't be overcome with a basic model and I don't think anyone is seriously arguing otherwise. For instance LLMs are pretty much stateless without their context window. If you treat the raw generated output as the first and final res…

Has anyone done the thing, and achieved some interesting results? Seems a pretty obvious thing to try, but I never heard of anything like it.

I noticed some examples from anthropic's golden-gate-claude paper had responses starting with for the inverse effect. Suppressing the output to the end of the paragraph would be an easy post processing operation.

It's probably better to have implicitly closed tags rather than requiring a close tag. It would be quite easy for a LLM to miss a close tag and be off in a dreamland.

Possibly addressing comments to the user or itself might allow for considering multiple streams of thought simultaneously. IRC logs would be decent training data for it to figure out many voice multi-conversations (maybe)

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#67

I am highly skeptical of LLMs as a mechanism to achieve AGI, but I also find this paper fairly unconvincing, bordering on tautological. I feel similarly about this as to what I've read of Chalmers - I agree with pretty much all of the conclusions, but I don't feel like the text would convince me of those conclusions if I disagreed; it's more like it's showing me ways of explaining or illustrating what I already belie…

Isn't any formal "proof" or "reasoning" that shows that something cannot be AGI inherently flawed, because we have a hard time formally describing what AGI is anyway.

Like your argument: embodiment is missing in LLMs, but is it needed for AGI? Nobody knows.

I feel we first have to do a better job defining the basics of intelligence, we can then define what it means to be an AGI, and only then can we prove that something is, or is not, AGI.

It seems that we skipped step 1 because its too hard, and jumped straight to step 3.

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#68

I am highly skeptical of LLMs as a mechanism to achieve AGI, but I also find this paper fairly unconvincing, bordering on tautological. I feel similarly about this as to what I've read of Chalmers - I agree with pretty much all of the conclusions, but I don't feel like the text would convince me of those conclusions if I disagreed; it's more like it's showing me ways of explaining or illustrating what I already belie…

Isn't any formal "proof" or "reasoning" that shows that something cannot be AGI inherently flawed, because we have a hard time formally describing what AGI is anyway. Like your argument: embodiment is missing in LLMs, but is it needed for AGI? Nobody knows. I feel we first have to do a better job defining the basics of intelligence, we can then define what it means to be an AGI, and only then can we prove that someth…

Yep, this is a big part of it. Intelligence and consciousness are barely understood beyond "I'll know it when I see it", which doesn't work for things you can't see - and in the case of consciousness, most definitions are explicitly based on concepts that are not only invisible but ineffable. And then we have no solid idea whether these things we can't really define, detect, or explain are intrinsically linked to each other or have a causal relationship in either direction. Almost any definition you pick is going to lead to some unsatisfying conclusions vis a vis non-human animals or "obviously not intelligent" forms of machine learning.

It's a real mess.

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#69
post #28
post #21

Earlier quoted context omitted.

> Future models will not be able to do those things if they are the same as the current ones I think a lot of people disagree with this. People think if we just keep adding parameters and data, magic will happen. That’s kind of what happened with ChatGPT after all.

I'm not so sure that view is very widespread amongst people familiar with how LLMs work. Certainly they become more capable with parameters and data, but there are fundamental things that can't be overcome with a basic model and I don't think anyone is seriously arguing otherwise. For instance LLMs are pretty much stateless without their context window. If you treat the raw generated output as the first and final res…

The problem with is that you need the internal monologue to not be subject to training loss, otherwise the internal monologue is restricted to the training distribution.

Something people don't seem to grasp is that the training data mostly doesn't contain any reasoning. Nobody has published brain activity recordings on the internet, only text written in human language.

People see information, process it internally in their own head which is not subject to any outside authority and then serialize the answer to human language, which is subject to outside authorities.

Think of the inverse. What if school teachers could read the thoughts of their students and punish any student that thinks the wrong thoughts. You would expect the intelligence of the class to rapidly decline.

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#70
post #30

Earlier quoted context omitted.

The question I gave is a literal textbook example of abstract reasoning. LLMs are just very advanced text prediction, but they are also provably capable of abstract reasoning. If you think that those statements are contradictory, I would encourage you to read up on the Bayesian hypotheses in cognitive science - it is highly plausible that our brains are also just very advanced prediction models.

You're quite right that LLMs can seemingly do some abstract reasoning problems, but I would not say they aren't in the training data. Sure, the exact form using the made up word gronk might not be in the training data, but the general form of that reasoning problem definitely exists, quite frequently in fact.

Have you seen this?

``` You will be given a name of an object (such as Car, Chair, Elephant) and a letter in the alphabet. Your goal is to first produce a 1-line description of how that object can be combined with the letter in an image (for example, for an elephant and the letter J, the trunk of the elephant can have a J shape, and for the letter A and a house, the house can have an A shape with the upper triangle of the A being the roof). Following the short description, please create SVG code to produce this (in the SVG use shapes like ellipses, triangles etc and polygons but try to defer from using quadratic curves). ```

``` Round 5: A car and the letter E. Description: The car has an E shape on its front bumper, with the horizontal lines of the E being lights and the vertical line being the license plate. ```

Image generated here: https://imgur.com/a/Ia4Q2h3

How does it "just" predict the letter E could be used in such a way to draw a car? How does it just text predict working SVG code that draws the car made out of basic shapes and the letter E?

I don't know how anyone could suggest there are no conceptual models embedded in there.

Post reply on HN