Live data from Hacker News

Telling GPT-4 you're scared or under pressure improves performance

aimodels.substack.com

221–230 of 255 posts

Re: Telling GPT-4 you're scared or under pressure improves performance

#221

> The implications are clear: incorporating emotional cues can lead to more effective and responsive AI applications. I think it's important to remember these models were trained on human interaction, and are in many ways a mirror for us to better understand human interaction. I don't think many people would be surprised if you said emotion can be used to better communicate with a human, but it is interesting to see…

This correlates with anecdata I've seen claiming that being polite with LLMs (using "please", "thank you", etc.) also improves performance. All in all, I like this. Both for what it implies about humans interacting with other humans, and for shaping norms about how best to to interact with other entities that show signs of intelligence.

I really dislike it.

It's a tool. It should execute instructions, not demand pleasantries.

Re: Telling GPT-4 you're scared or under pressure improves performance

#222

Earlier quoted context omitted.

Sure it does. It perfectly deliniates it. LLMs are not: sensitive to causal structure, dynamically adapting to environmental changes, growing, developing sensory-motor capacities, they are not with us in our environment, they are not: expressing desires, preferences, intentions, beliefs, motivations, etc. And so on. To say, "they just predict the next word" is literally to say that all apparent functions of an LLM ar…

You didn't really address what og_kalu brought up. Which is that, it's possible that the model learns human like thinking, because that's the best way to accurately predict the human response itself. I generally agree with you, but still, I do think that this is the current question. What it is the model is learning that it then uses for predictions? Because you're assuming it's learning some purely token correlation…

I wasnt getting the sense it was worthwhile to engage, as my views werent being accurately understood.

By I can address this. The meaning of words is, roughly, states of the world. If I say, "pass me the salt" that is satisfied if you, in fact, pass me the salt. If I say, "that tree is green" this is true if that tree which we are both talking about has the property of causing a perceptual state "seeming green" in both of us. And so on.

The distribution of text has nothing to do with the meaning of words. Rather, we language users, for convenience, arrange words in orders that are related to their actual meaning. It is our ordering, for communicative convenience, that makes 'replaying the distributions of text' back to us apparently successful.

But, strictly, there isnt anything for the LLM to learn as far as meaning goes. It simply doesnt have the data to acquire the meaning of words. Not untill it can pass salt can it ever mean to say, "pass me the salt" and so on.

For any given sentence consider what capacities an agent would have to have in order to mean it. Consider, "I liked that film!", "I wish I was in france", "I believe the car outside is a BMW", and so on. These concern internal capacities (aethetic judgement, imagination, propositional attitudes, representational attitudes, etc.) and their orientation to an external world (the film, france, the car, ... you, me, etc.). Capacities profoundly absent here.

The methodological premise of your question is that if a system has text inputs and outputs that match 'human competence' restricted to the domain of text inputs and outputs -- then we should assume similar capacities.

But this is trivial to disprove. Assume there exists a dictionary from all prompts to all answers, then this dictionary has human-level 'competence'. But a dictionary lookup does not employ any human capacities: no imagination, no reasoning, etc.

So we cannot do this, really quite dumb thing, of saying "well i'm fooled by these prompts and their answers" and thereby impart, in total ignorance, capacities to a system. This, really seriously, is pseudoscience.

Science would be to start with a theory of these capacities, ie., of imagination, belief, represtational states, attitudes to the world, and so on -- then determine empirical tests for their presence in a system, and then determine if LLMs could even have them.

If you do this, however, you immediately rule out all systems which merely map text to text. We do not determine, say, whether an animal can imagine an alternative possibility by feeding it some text input.

The very form that "AI" here takes already precludes being intelligent. Intelligene, as a natural phenomenon, is not an implementation of a function from text to text. This incredibly restricted domain is indeed a clue that it's a trick.

Saying, "you can only use text" is just like the magician saying, "please, stay seated" (the trick only works if you dont move).

There is nothing an LLM could do to meet any plausible empirical theory of intelligence. If you gave me 100% human competence on all prompts, that's really entirely irrelevant.

Prompts are not a test of any capacity. The success criterion of AI engineers, that of 'accuracy' is an engineering metric, not a scientific one. It's pseudoscience to say that covering some (Q, A) to 100% implies the system can imagine, say, or anything else.

This is just confused thinking. Bugs bunny can speak as well as he likes, that does not mean he's witty -- he doesnt exist.

The turing test, as well as all mathematical criteria of domain-covering accuracy, are tests of how well we have fooled users. They arent science.

Re: Telling GPT-4 you're scared or under pressure improves performance

#223
post #51

Earlier quoted context omitted.

It is not accurate to say that an LLM like ChatGPT predicts anything. It is trained to maximize a score function, so it is more like trying to win a game where the moves are word choices.

The game is predicting the next word a person would write.

That’s not true and that was the entire point of my comment. It’s like playing chess, you’re not predicting your next best move, you’re searching for it.

Re: Telling GPT-4 you're scared or under pressure improves performance

#224

Earlier quoted context omitted.

Your comment literally reads: "LLMs predict words. Any semantic validity is a side effect of enough training data reinforcing the close correlation of those tokens." How am I supposed to interpret this any other way? If your claim is that LLMs currently do not possess the same generalization ability as humans, then no one here would disagree with you. But you went way further by claiming that only close correlations…

I think LLMs are a good start. I am certain they lack a world model, the kind you and me use. This is, to me, a fact. I think that eventually we will bridge these gaps. I also work on implementations that smash into the limits I am describing. I am not the only one. I have scrupulously avoided calling it hallucinations, but these are the litmus test where the claims fail. The failures are not a case of not knowing sp…

The issue I have with your comments is that you make some reasonable points, and then immediately over-extrapolate these points unreasonably.

>I think LLMs are a good start. I am certain they lack a world model, the kind you and me use.

See, I agree here, with emphasis on "the kind you and me use". Yes, we have a greater capability to generalize than current LLMs, that is clear.

> The failures are not a case of not knowing specific nouns, they are a generalization failure that a world model would prevent.

And then you say something like this, which is obviously wrong. No, a world model wouldn't prevent generalization failures, a perfect all-encompassing world model would. Humans experience generalization failures as well, otherwise every athlete in one sport would automatically be an expert in every other discipline or every mathematician would also be a Grand Master in chess. LLMs necessarily need a world model to generate well-formed text that isn't in their training corpus, something they are obviously capable of, it's just an imperfect world model. Ours is also imperfect, but far less so than that of LLMs.

> I have linked a paper...

Except that paper is completely irrelevant to the argument you're making here. It is a useful insight into the limitations of simple metrics, but definitely does not extend to any claim of model performance, because they too use a simple metric as an replacement, even though clear qualitative differences are observed between model iterations.

Let me put it this way: Imagine I create a series of chess AIs, with each iteration better than the last. If I then show you a chart demonstrating that the ELO of my models increases linearly, would you say that my models' abilities increase linearly as well? No, obviously not, because my model needs far less strategy and complexity to go from ELO 1000 to 1100 than it needs to go from 2700 to 2800. I.e the difficulty doesn't scale linearly, and a linear increase on this nonlinear space is therefore also not really linear. Unless you believe the difficulty of accurately predicting text scales linearly, then this applies to LLMs as well.

> If your model decides that a rose by any other name doesn’t smell just as sweet, then your model is fundamentally not seeing roses.

Except that this is the entire value proposition of LLMs. They can, in the average case, actually represent concepts by the complex interplay of adjacent concepts. The entire reason why they are so impressive is that the nuances of reality are grasped and that even a noisy example of a concept can be correctly classified. Give a LLM a description that is largely incorrect and mislabeled, and chances are it gets it anyway. LLMs being unable to generalize over some concepts has as much to do with fundamental limitations as me being unable to correctly classify the shredded remains of a flower variety that I have seen once in my life has to do with me being stupid.

> Look, you can argue with me or you can try it out. Push the system, see how far it can go

I have done just that for the last 6 months and have seen nothing to contradict what I've said here.

Re: Telling GPT-4 you're scared or under pressure improves performance

#227

Earlier quoted context omitted.

most formulas, including that one, are neural networks (such is the absurdity of the term) -- so it is trivially learnable The issue is that to prepare the dataset from which that formula is learnt requires already knowing it. This is the triviality of applications of universal function approximators to science -- empirical data modelling isnt new, and neural networks are just one example of it; not all that special.…

> The issue is that to prepare the dataset from which that formula is learnt requires already knowing it. Why is that? Are you saying that universal function approximators can't invent higher level models? Isn't that what happens in the hidden layers when an neural network is trying to predict things that require those higher level models?

No, consider how you would collect a dataset to show F=GMm/r^2

you'd need to measure the force between planets, the distance between them, etc. no such direct measurements have taken place to my knowledge, certainly not at newton's time

All "universal fn approximator" means is that given any dataset which is sampled from a function, you can recover that function with enough data points. It does not mean, eg., that given a 2D dataset you can derive a 3D function -- you cannot, this is impossible.

So you need to be sampling from F=GMm/r^2 to find that function. So you need to know the answer before you begin. Fn approximators are only useful for empirical refinements to existing knowledge.

In order to construct such datasets you need to do science; hence experiments, etc.

Re: Telling GPT-4 you're scared or under pressure improves performance

#228

Earlier quoted context omitted.

>Since it's incapable of actually summarising financial data It's not, though. It is in fact able to summarize financial data, just as it's able to write code and diagnose a medical condition. It makes mistakes, yes, even grave ones, much more so than experts in those fields would.

It isnt making mistakes ... its never actually doing it. Do you see a difference between the process of adding numbers and dividing by their count (taking a mean) and emitting numeric tokens which are most probable for a given input? The former is called "taking a mean" the latter isnt . This system never engages in any method to summarise financial data. It's method is always the same: to emit tokens most probable g…

To add to your point: try asking ChatGPT to do basic arithmetic on numbers it hasn't seen before. You'll see just how good it is at computation.

Re: Telling GPT-4 you're scared or under pressure improves performance

#229

Earlier quoted context omitted.

You didn't really address what og_kalu brought up. Which is that, it's possible that the model learns human like thinking, because that's the best way to accurately predict the human response itself. I generally agree with you, but still, I do think that this is the current question. What it is the model is learning that it then uses for predictions? Because you're assuming it's learning some purely token correlation…

I wasnt getting the sense it was worthwhile to engage, as my views werent being accurately understood. By I can address this. The meaning of words is, roughly, states of the world. If I say, "pass me the salt" that is satisfied if you, in fact, pass me the salt. If I say, "that tree is green" this is true if that tree which we are both talking about has the property of causing a perceptual state "seeming green" in bo…

Why would an LLM require salt?

Re: Telling GPT-4 you're scared or under pressure improves performance

#230

Earlier quoted context omitted.

What if humans’ responses are merely probabilistically consistent with a history of sensory experiences? Would this change the significance of human emotions vs apparent emergent emotional responses from LLMs?

emotions regulate motivation, desire, action, behaviour etc. to be angry is for your sensory-motor system to be primed for aggression; it's for your cognitive systems to be narrowed and focused on analysing high-threat parts of your environment; it is for your memory-formulation to be modulated towards threat recollection etc. Sure, if an LLM's prompt "be angry" causes it to adopt a threat stance to its environment,…

Why should I discount a theory just to protect my ego? We've read countless stories about science only progressing when those with big egos die. It would only seem logical that eventually it will come for my own.
Post reply on HN