Live data from Hacker News

Large models of what? Mistaking engineering achievements for linguistic agency

arxiv.org

101–110 of 162 posts

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#101
post #86

Earlier quoted context omitted.

Put aside consciousness or hype or investment. Look at the results; LLMs are well beyond old-school search in many ways. Sure, they are flawed in someways. Previous paradigms for search, were also flawed in their own ways. Look at the arc of NLP. Large language models fit the pattern. One could even say that their development (next token prediction with a powerful function approximator) is obvious in hindsight.

Honestly I don't disagree, I just think that humans tend to anthropomorphize to such a high extent that there is a fair bit of hyperbole promoting LLMs as more than they are. It's my opinion that the big flaws LLMs currently present aren't going to be overcome by scaling alone.

Scaling existing architectures (inference I mean) will probably help a lot. Combine that with better training and hybrid architectures, and I personally expect to see continued improvement.

However, given the hype cycle, combined with broad levels of ignorance of how LLMs work, it is an open question if even amazing progress will impress people anymore.

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#102
post #72

Earlier quoted context omitted.

Transformer models have been shown to spontaneously form internal, predictive models of their input spaces. This is one of the most pervasive misunderstandings about LLMs (and other transformers) around. It is of course also true that the quality of these internal models depends a lot on the kind of task it is trained on. A GPT must be able to reproduce a huge swathe of human output, so the internal models it picks o…

> can provide links if you're interested Please do :)

Here's the paper on OthelloGPT's internal models I mentioned: https://arxiv.org/abs/2309.00941

The references in that paper are also good reading!

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#103
post #94

Earlier quoted context omitted.

I can't disagree more. Or maybe I actually agree. Because it's not easy to tell whether something is flying. Definitions like that fall apart every time we encounter something out of the ordinary. If you take the criterion of "there's no discussion about it", then you're limiting the definition to that which is familiar, not that which is interesting. Is an ekranoplan flying? Is an orbiting spaceship flying? Is a hov…

"It's not flying, it's falling... with style"

I always fall with style and I always do it on purpose :|

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#104
post #84

The authors of this paper are just another instance of the AI hype being used by people who have no connection to it, to attract some kind of attention. "Here is what we think about this current hot topic; please read our stuff and cite generously ..." > Language completeness assumes that a distinct and complete thing such as `a natural language' exists, the essential characteristics of which can be effectively and c…

Please argue on the merits and substance. I’m less interested in speculation on the authors’ motivations.

The parent comment I responded to is speculative and does not argue on the merits. We can do better here.

Are there people who ride the hype wave of AI? Sure.

But how can you tell from where you sit? How do you come to such a judgment? Are you being thoughtful and rational?

Have you considered an alternative explanation? I think the odds are much greater that the authors’ academic roots/training is at odds with what you think is productive. (This is what I think, BTW. I found the paper to be a waste of my time. Perhaps others can get value from it?)

But I don’t pretend to know the authors’ motivations, nor will I cast aspersions on them.

When one casts shade on a person like the comment above did, one invites and deserves this level of criticism.

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#105

I am highly skeptical of LLMs as a mechanism to achieve AGI, but I also find this paper fairly unconvincing, bordering on tautological. I feel similarly about this as to what I've read of Chalmers - I agree with pretty much all of the conclusions, but I don't feel like the text would convince me of those conclusions if I disagreed; it's more like it's showing me ways of explaining or illustrating what I already belie…

There are many finite problems that absolutely do not admit finite solutions. Full stop. I think the deeper point of the paper is that you simply cannot generate an intelligent entity by just looking at recorded language. You can create a dictionary, and a map - but one must not mistake this map for the territory.

The human brain is a finite solution, so we already have an existence proof. That means a lot for our confidence in the solvability of this kind of problem.

It is also not universally impossible to reconstruct a function of finite complexity from only samples of its inputs and outputs. It is sometimes possible to draw a map that is an exact replica of the territory.

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#106

I am highly skeptical of LLMs as a mechanism to achieve AGI, but I also find this paper fairly unconvincing, bordering on tautological. I feel similarly about this as to what I've read of Chalmers - I agree with pretty much all of the conclusions, but I don't feel like the text would convince me of those conclusions if I disagreed; it's more like it's showing me ways of explaining or illustrating what I already belie…

To me LLMs seem to most closely resemble the regions of the brain used for converting speech to abstract thought and vice-versa, because LLMs are very good at generating natural language and knowing the flow of speech. An LLM is similar to if you took the the Wernicke's and Broca's Areas and stuck a regression between them. The problem is that the regression in the middle is just a brute force of the entire world's knowledge instead of a real thought.

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#107

Earlier quoted context omitted.

I don't even know where to begin to address your confusion. Without computability theory there are no computers, no operating systems, no networks, no compilers, and no high level frameworks for "AI".

Well, if you want to address my "confusion" then pick something and start there =) That is patently false - most of those things are firmly in the realm of engineering, especially these days. Mathematics is good for grounding intuition though. But why is this relevant to the OP?

There is no reason to do any of that because according to your own logic AI can do all of it. You really should sit down and ponder what exactly you get out of equating Turing machines with human intelligence.

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#108

Earlier quoted context omitted.

Well, if you want to address my "confusion" then pick something and start there =) That is patently false - most of those things are firmly in the realm of engineering, especially these days. Mathematics is good for grounding intuition though. But why is this relevant to the OP?

There is no reason to do any of that because according to your own logic AI can do all of it. You really should sit down and ponder what exactly you get out of equating Turing machines with human intelligence.

Sorry, I edited my reply because I decided going down that rabbit hole wasn't worth it. Didn't expect you to reply immediately.

I'm not equating anything here, just pointing out that the fact that AI runs in software isn't a knockdown argument against anything. And computability theory certainly has nothing useful to say in that regard.

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#109

Earlier quoted context omitted.

I can't disagree more. Or maybe I actually agree. Because it's not easy to tell whether something is flying. Definitions like that fall apart every time we encounter something out of the ordinary. If you take the criterion of "there's no discussion about it", then you're limiting the definition to that which is familiar, not that which is interesting. Is an ekranoplan flying? Is an orbiting spaceship flying? Is a hov…

There are going to be gray areas of course, but the point I'm making is that if it's hard to argue something isn't flying (respectively, intelligent) then it's probably flying (resp. intelligent). If it's hard to tell then it's probably not. I'm suggesting that intelligence, like flying, should be very immediately obvious. For example, you can't miss the fact that a five-year old child is intelligent and you can't mi…

> when something is intelligent then it should leave us no doubt that it is.

I strongly disagree. There are many reasons we might not recognize its intelligence, such as:

- it operates on a different timescale than we do.

- it operates at a different size scale than we do.

- we don't understand its language, its methods, or its goals

- Cartesian-like ideological blindness ("only humans have experience, all other things are automata, no matter how much they seem otherwise")

Throughout human history, certain people have even managed to doubt the intelligence of other groups of humans.

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#110

I am highly skeptical of LLMs as a mechanism to achieve AGI, but I also find this paper fairly unconvincing, bordering on tautological. I feel similarly about this as to what I've read of Chalmers - I agree with pretty much all of the conclusions, but I don't feel like the text would convince me of those conclusions if I disagreed; it's more like it's showing me ways of explaining or illustrating what I already belie…

> LLMs do not have corporeal experience. But it's not obvious that this means that they cannot, a priori, have an "internal" concept of reality, or that it's impossible to gain such an understanding from text. I would argue it is (obviously) impossible the way the current implementation of models work. How could a system which produces a single next word based upon a likelihood and and a parameter called a "temperatu…

> How could a system which produces a single next word based upon a likelihood and and a parameter called a "temperature" have a conceptual model underpinning it? Even theoretically?

You're limiting your view of their capabilities on the output format.

> Not so with LLMs!! Generative LLMs do not have a prior concept available before they start emitting text.

How do you establish that? What do you think of othellogpt? That seems to form an internal world model.

> That the "temperature" can chaotically change the output as the tokens proceed

Changing the temperature forcibly makes the model pick words it thinks fit worse. Of course it changes the output. It's like an improv game with someone shouting "CHANGE!".

Let's make two tiny changes.

One, let's tell a model to use the format

askjdhas as the voice in their head, and blah for the output.

Second, let's remove temperature and keep it at 0 so we're not playing a game where we force them to choose different words.

Now what remains of the argument?

Post reply on HN