Live data from Hacker News

Telling GPT-4 you're scared or under pressure improves performance

aimodels.substack.com

241–250 of 255 posts

Re: Telling GPT-4 you're scared or under pressure improves performance

#241
post #214

Earlier quoted context omitted.

You are correct, but also why would you say this is an "engineering trick"? The interesting part about LLMs is what they are capable of doing and how they are constructed. It's not a great argument to say that they don't have motor functions (!?!?). It shows the incredible ability of deep learning to discover and operate with high-level concepts.

Because he doesn't believe a plane can fly without flapping wings and feathers. That's all this boils down to. To him, plane flight must also be an "engineering trick" or in other words, his idea of flight has been divorced of any real meaning.

In the case of flight we're interested in a function: transportation. The plane performs that function so we call it 'flight'.

Were we interested in navigating the world as a flying animal does: in flocks, social navigation, hunting, etc. then indeed, planes would not count. Planes, in this sense, do not fly. Planes, in this sense, are a trick.

We already have 'functional intelligence', we've had it for thousands of years and perfected it in the 20th C with the electronic computer. Any system of automation we build is 'functionally intelligent', including the plane.

The problem is were not interested in this form. When Commander Data was written, with "I, Robot" and the like, the authors were writing People. They were writing animals. They imagined flying birds, just made of metal.

This is a (nomenological) impossibility, just as impossible as making an aeroplane to fly like a bird.

The kind of intelligence which matters to us is animal intelligence. The kind which matters to engineers, of course, is purely functional: more trinkets to sell.

But you cannot glue these together and call them a person, nor pronounce that what matters to you is the same as what matters to everyone else. This is a delusion, or a lie, and largely a mixture of both. A series of lies told to the public to maintain a popular delusion, one which is profoundly dangerous.

A creature in the world with us, talking to us and meaning what it says, judging situations we are in, advising us based on our needs and its understanding of them -- etc. is a creature whose mode of operation and mode of life is alike our own.

Lying that LLMs are this, saying that 'planes flock together in the sky', is dangerous. Users of these systems adopt a schizophrenic disposition to them, and rely on them -- this reliance, based on a lie, is dangerous to them. These systems have no such capacities. They generate text according to what is, on average, best given a historical corpus.

They do not imagine, reflect, emote. They do not move, sense, or coordinate. They have no intention to speak, and cannot mean what they say. They are not here with us. 'They' are not a 'they' at all -- rather, cleverly constructed tea-leaves based on a shredded recording of everything ever written.

In the matters of intelligence, we want an animal -- we do not want a calculator. This is a solved problem. And it is a dangerous thing to tell people their calculator's advice has considered their interests -- what nonesense.

Re: Telling GPT-4 you're scared or under pressure improves performance

#242
post #49

Earlier quoted context omitted.

This wasn't done deliberately. It's something the model picked up from it's training dataset. I'm sure the creators don't want it to be like this, but cleaning trillions of tokens of all examples is long and hard work. The alternative would be injecting these sentences into your prompts, which probably nobody really wants to happen.

I know it's not deliberate, but since they know this improves results, shouldn't they modify the service so that it takes advantage of this sort of improvement automatically?

But how would they do that? I don't want my queries to be rewritten, or new parts to be injected, because it would limit my ability to actually write the exact query I need.

Re: Telling GPT-4 you're scared or under pressure improves performance

#243

Earlier quoted context omitted.

I think LLMs are a good start. I am certain they lack a world model, the kind you and me use. This is, to me, a fact. I think that eventually we will bridge these gaps. I also work on implementations that smash into the limits I am describing. I am not the only one. I have scrupulously avoided calling it hallucinations, but these are the litmus test where the claims fail. The failures are not a case of not knowing sp…

The issue I have with your comments is that you make some reasonable points, and then immediately over-extrapolate these points unreasonably. >I think LLMs are a good start. I am certain they lack a world model, the kind you and me use. See, I agree here, with emphasis on "the kind you and me use". Yes, we have a greater capability to generalize than current LLMs, that is clear. > The failures are not a case of not k…

I suspect we are getting into an issue of degrees, potentially due to differences in how you and I have been applying LLMs.

For example you said that in the average case they actually represent concepts by the complex interplay of adjacent concepts - I would agree. ChatGPT can pass the bar, it can pass medical exams etc. I would also point out that the work in that sentence is being done by the term “average case”.

Let’s assume our experiences diverge at this point. At the start of the year, I started tinkering, then actively trying to push LLMs to failure, in order to understand the limits of what could be achieved.

After creating several tools/experiments you end up having to deal with Hallucinations, and this is where my stance likely diverged from yours.

Two different studies showed generated content was only ~50% and ~40% supported by provided citations.

One out of 4 of my summarization tests was spectacularly fabricated.

I had bad performance on even classification tasks - and OpenAI engineers described this same failure at a conference. I am a recovering non-coder, so you dont have to take my word for it.

At work, I need processes that are more than ~97.x% accurate, otherwise they are poor replacements for the human in the loop ones already in place.

Average case performance suggested the ability to actively plan, to actively assess situations. However hallucinations overrode those capabilities. LLMs will actively imagine functions, teams, or processes that dont exist.

Eventually, it became clear that LLM hallucination is far too anthropomoprhized a word. LLMs are always “hallucinating” - it’s only humans who have an issue with the output.

If I have understand you correctly, semantic inaccuracy to you is simply a failure of not having enough related concepts.

I wish I could remember the exact examples that made me realize there is no world view at play at all, I could simply share those.

Instead, can you describe how you are getting acceptable performance from your LLMs? Maybe experience and use cases will be enough to bridge the gap.

Re: Telling GPT-4 you're scared or under pressure improves performance

#244

Am I the only one who feels bad for asking ChatGPT a "dumb" question that I know I should know, or not saying thank you when it gives me an answer? No? I'm just a weirdo? Okay. I have to push back really hard against my proclivity to humanize it, to the point where I probably don't use it as much as I should, just because I don't want to deal with the psychic stress of reminding myself that it's not a living entity.

> Am I the only one who feels bad for asking ChatGPT a "dumb" question that I know I should know, or not saying thank you when it gives me an answer? Have you tried asking ChatGPT? Jokes aside, I feel it sometimes as well. I wonder if it's more because I say "thank you" more as a reflex than an actual feeling of gratitude for most of my human interactions.

> "Have you tried asking ChatGPT?"

I am so sick of that "joke" already, and it feels like it is there to stay just like we had "just ask Google" for the last 25 years. It is so unimaginative.

Re: Telling GPT-4 you're scared or under pressure improves performance

#245

Earlier quoted context omitted.

Sure it does. It perfectly deliniates it. LLMs are not: sensitive to causal structure, dynamically adapting to environmental changes, growing, developing sensory-motor capacities, they are not with us in our environment, they are not: expressing desires, preferences, intentions, beliefs, motivations, etc. And so on. To say, "they just predict the next word" is literally to say that all apparent functions of an LLM ar…

The problem is the word "just". Saying they just predict the next word in a sequence is where the statement jumps from being a straightforward factual scientific claim, to one that contains an opinion. After all, if it just predicts the next word, the unspoken implication is it can't be very that good. It's a shallow dismissal of a collosal amount of work. Do you consider "typist" an accurate description of your job?

No, because I have reasons to type and it is those reasons which express what I am doing.

An LLM has no reason to be doing anything; it is not responsive to reasons.

If I say, "i like what you're wearing" i may: be flitring, expressing my taste, being enouraging, etc -- perhaps all at once.

It is precisely all these reasons for action which LLMs lack. They generate text on the occasion of a prompt, not for any reason (in this sense) at all. So literally: they do not act.

An LLM is more like a river than a person. The flow of electrons which brings about a response to a prompt is a (very narrowly) deterministic function of a historical corpus of text.

Whereas a person is a narrowly non-deterministic, or very broadly deterministic, function of their experiences and capacities. People grow in their environments, and in growing, acquire novel dispositions which give them reasons for acting.

The word "just" here is very important. They are, very much, just generating text.

There really isnt any significant achivement here at all. Big tech companies stole decade's worth of our electronic data --- comments, books, forums we created to share with each other -- and ran it through many-mil-$ hardware costing many-mil-$ electricity. They ran it through a fundamentally simple algorithm.

All the achivement here is ours as a species communicating digitally and recording our lives. I regard OpenAI, et al. as profoundly parasitical on this. Replaying ourselves back to us, and claiming "ChatGPT" as an author.

This scam-framing hoodwinks investors, and the public, into ever higher valuations based on ever more ridiculous hagiography. There is a tool here, and it's value comes from us

Re: Telling GPT-4 you're scared or under pressure improves performance

#246

Earlier quoted context omitted.

> it is now quite well established that GPT-4 has impressive out-of-sample performance Err... I can show this is false, kinda trivially. People who engage in prompt-confirmation-bias aren't aware of what the in-sample is. It's basically everything ever digitised: you can ask it for the first paragraph of every dickens novel, to what the average petal length of an iris flower is -- etc. How are you measuring the in-sa…

You can make it invent a new language: https://maximumeffort.substack.com/p/i-taught-chatgpt-to-inv... I am sure you will continue to argue that this is still in line with everything-thats-ever-written prediction but my opinion is that at that point, it's a meaningless distinction. The human brain is also just a machine.

But “everything ever digitised” includes a tonne of linguistics information - it’s still in sample.

Re: Telling GPT-4 you're scared or under pressure improves performance

#247

Earlier quoted context omitted.

Depends. Models are matrices of floats and so there's little chance an umbrella-term like "stochastic parrot" will never not stick, even when they already show signs of syntactic, semantic world-building capability ( https://www.arxiv-vanity.com/papers/2206.07682/ ). If you are like me (and them: https://archive.is/cZi83 ) and deem instruction following , chain-of-thought prompting , computational properties of LLMs…

Okay so just to confirm that section doesn't actually tell us anything about this and in fact this is all based on your own understanding of the mechanisms involved.

My reading of the papers is, given enough scale, modality, and memory; there are chances (perhaps newer and different) models will be able to "generalize" our world. Also: https://archive.is/3yyZZ / https://twitter.com/QuanquanGu/status/1721394508146057597 | And: https://archive.is/bW2tS / https://twitter.com/mansiege/status/1680985267262619648

Re: Telling GPT-4 you're scared or under pressure improves performance

#248
post #221

Earlier quoted context omitted.

This correlates with anecdata I've seen claiming that being polite with LLMs (using "please", "thank you", etc.) also improves performance. All in all, I like this. Both for what it implies about humans interacting with other humans, and for shaping norms about how best to to interact with other entities that show signs of intelligence.

I really dislike it. It's a tool. It should execute instructions, not demand pleasantries.

A slave in the Roman empire/United States/Nazi concentration camps etc etc were also perceived as a tool. Do they deserve pleasantries? (no implication about morality of slavery but I intentionally selected examples where enslaved people were dehumanized)

Re: Telling GPT-4 you're scared or under pressure improves performance

#249

Earlier quoted context omitted.

> I have to push back really hard against this stuff Based on ChatGPT's answers, OpenAI think you should be saying thanks to it, because it helps you have a more natural conversation, which encourages you as the user to send more natural/productive prompts when asking real questions. As for asking it dumb questions, how often do you use Google as a spellcheck? I wouldn't consider this any different.

>how often do you use Google as a spellcheck? // When Word tells me a word is wrong, but I'm almost certain it's right, and it turns out not to be in the Microsoft dictionary somehow. I mean, I have an extensive vocabulary, but not more extensive than a decent British-English dictionary.

Spell check is crazy inconsistent and poorly implemented. I'm constantly running across words the browser says is misspelled but isn't. I know you can replace the built in spell check with a better dictionary, but trying to keep that in sync across all my devices just keeps me going back to Google as spellcheck.

Re: Telling GPT-4 you're scared or under pressure improves performance

#250

I think this is the entry point needed to get peoples attention and explain: LLMs aren’t people, and emergent properties are being over extended. If LLMs are showing “better” performance when there are tokens that humans read as emotionally salient - Then the underlying text it’s trained on shows humans give better answers when emotionally salient context is provided. LLMs predict words. Any semantic validity is a si…

Check out task-oriented manifolds that emerge from training.

https://www.cell.com/neuron/pdf/S0896-6273(19)31044-X.pdf

Post reply on HN