Live data from Hacker News

Large models of what? Mistaking engineering achievements for linguistic agency

arxiv.org

41–50 of 162 posts

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#41
post #6

Full paper: [1]. Not much new here. The basic criticism is that LLMs are not embodied; they have no interaction with the real world. The same criticism can be applied to most office work. Useful insight: "We (humans) are always doing more than one thing." This is in the sense of language output having goals for the speaker, not just delivering information. This is related to the problem of LLMs losing the thread of a…

The optimization process adjusts the weights of a computational graph until the numeric outputs align with some baseline statistics of a large data set. There is no "punishment" or "reward", gradient descent isn't even necessary as there are methods for modifying the weights in other ways and the optimization still converges to a desired distribution which people claim is "intelligent". The converse is that people ar…

If you really want to phrase it that way, organisms like us are "just" distributions of genes that have been pushed this way and that by natural selection until they converged to something we consider intelligent (humans).

It's pretty clear that these optimisation processes lead to emergent behaviour, both in ML and in the natural sciences. Computability theory isn't really relevant here.

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#42

I am highly skeptical of LLMs as a mechanism to achieve AGI, but I also find this paper fairly unconvincing, bordering on tautological. I feel similarly about this as to what I've read of Chalmers - I agree with pretty much all of the conclusions, but I don't feel like the text would convince me of those conclusions if I disagreed; it's more like it's showing me ways of explaining or illustrating what I already belie…

> LLMs do not have corporeal experience. But it's not obvious that this means that they cannot, a priori, have an "internal" concept of reality, or that it's impossible to gain such an understanding from text. I would argue it is (obviously) impossible the way the current implementation of models work. How could a system which produces a single next word based upon a likelihood and and a parameter called a "temperatu…

Transformer models have been shown to spontaneously form internal, predictive models of their input spaces. This is one of the most pervasive misunderstandings about LLMs (and other transformers) around. It is of course also true that the quality of these internal models depends a lot on the kind of task it is trained on. A GPT must be able to reproduce a huge swathe of human output, so the internal models it picks out would be those that are the most useful for that task, and might not include models of common mathematical tasks, for instance, unless they are common in the training set.

Have a look at the OthelloGPT papers (can provide links if you're interested). This is one of the reasons people are so interested in them!

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#43

The authors of this paper are just another instance of the AI hype being used by people who have no connection to it, to attract some kind of attention. "Here is what we think about this current hot topic; please read our stuff and cite generously ..." > Language completeness assumes that a distinct and complete thing such as `a natural language' exists, the essential characteristics of which can be effectively and c…

Babies have feedback and interaction with someone speaking to them. Would they learn to speak if you just dumped them in front of a TV and never spoke to them? I'm not sure.

But anyway I agree with you. This is just a confused HN comment in paper form.

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#44

I am highly skeptical of LLMs as a mechanism to achieve AGI, but I also find this paper fairly unconvincing, bordering on tautological. I feel similarly about this as to what I've read of Chalmers - I agree with pretty much all of the conclusions, but I don't feel like the text would convince me of those conclusions if I disagreed; it's more like it's showing me ways of explaining or illustrating what I already belie…

The crux of the video game analogy seems to be that when you go close to an object, the resolution starts blurring and the illusion gets broken, and there is a similar thing that happens with LLMs (as of today) as well. This is, so far, reasonable based on daily experience with these models.

The extension of that argument being made in the paper is that a model trained on language tokens spewed by humans is incapable of actually reaching that limit where this illusion will never breakdown in resolution. That also seems reasonable to me. They use the word "languaging" in verb form as opposed to "language" as a noun to express this.

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#45

The authors of this paper are just another instance of the AI hype being used by people who have no connection to it, to attract some kind of attention. "Here is what we think about this current hot topic; please read our stuff and cite generously ..." > Language completeness assumes that a distinct and complete thing such as `a natural language' exists, the essential characteristics of which can be effectively and c…

They are two researchers/assistant professors working with cognitive science, psychology, and trustworthy AI. The paper is peer reviewed and has been accepted for publication in the Journal of Language Sciences.

You should publish your critique of their research in that same journal.

P.s. if you find any grave mistakes, you can contact the editor in chief, who happens to be a linguist.

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#46

I am highly skeptical of LLMs as a mechanism to achieve AGI, but I also find this paper fairly unconvincing, bordering on tautological. I feel similarly about this as to what I've read of Chalmers - I agree with pretty much all of the conclusions, but I don't feel like the text would convince me of those conclusions if I disagreed; it's more like it's showing me ways of explaining or illustrating what I already belie…

> . I feel similarly about this as to what I've read of Chalmers - I agree with pretty much all of the conclusions, but I don't feel like the text would convince me of those conclusions if I disagreed;

my limited experience of reading Chalmers is that he doesn't actually present evidence - he goes on a meandering rant and then claims to have proved things that he didn't even cover. it was the most infuriating read of my life, I heavily annotated two chapters and then finally gave up and donated the book.

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#47

Earlier quoted context omitted.

The optimization process adjusts the weights of a computational graph until the numeric outputs align with some baseline statistics of a large data set. There is no "punishment" or "reward", gradient descent isn't even necessary as there are methods for modifying the weights in other ways and the optimization still converges to a desired distribution which people claim is "intelligent". The converse is that people ar…

If you really want to phrase it that way, organisms like us are "just" distributions of genes that have been pushed this way and that by natural selection until they converged to something we consider intelligent (humans). It's pretty clear that these optimisation processes lead to emergent behaviour, both in ML and in the natural sciences. Computability theory isn't really relevant here.

I don't even know where to begin to address your confusion. Without computability theory there are no computers, no operating systems, no networks, no compilers, and no high level frameworks for "AI".

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#48

The authors of this paper are just another instance of the AI hype being used by people who have no connection to it, to attract some kind of attention. "Here is what we think about this current hot topic; please read our stuff and cite generously ..." > Language completeness assumes that a distinct and complete thing such as `a natural language' exists, the essential characteristics of which can be effectively and c…

They are two researchers/assistant professors working with cognitive science, psychology, and trustworthy AI. The paper is peer reviewed and has been accepted for publication in the Journal of Language Sciences. You should publish your critique of their research in that same journal. P.s. if you find any grave mistakes, you can contact the editor in chief, who happens to be a linguist.

> You should publish your critique of their research in that same journal.

No thanks; that would be at least twice removed from Making Stuff.

(Once removed is writing about Making Stuff.)

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#49
Where I work, there's a somewhat haphazardly divided org structure, where my team has some responsibility to answer the executives demands for "use AI to help our core business". So we applied off-the-shelf models to extract structured context from mostly unstructured text - effectively a data engineering job - and thereby support analytics and create more dashboards for the execs to mull over.

Another team, with a similar role in a different part of the org has jumped (feet first) into optimizing large language models to turn them into agents, without consulting the business about whether they need such things. RAG, LoRA and all this optimization is well and good, but this engineering focus has found no actual application, expect wasting several million bucks hiring staff to do something nobody wants.

Post reply on HN