Live data from Hacker News

Telling GPT-4 you're scared or under pressure improves performance

aimodels.substack.com

91–100 of 255 posts

Re: Telling GPT-4 you're scared or under pressure improves performance

#91

Am I the only one who feels bad for asking ChatGPT a "dumb" question that I know I should know, or not saying thank you when it gives me an answer? No? I'm just a weirdo? Okay. I have to push back really hard against my proclivity to humanize it, to the point where I probably don't use it as much as I should, just because I don't want to deal with the psychic stress of reminding myself that it's not a living entity.

I have had to resist the temptation to thank ATMs before, and they don't model human interaction in a way which takes into account correlations between politeness and responsiveness...

As for silly questions, I'd be more worried I might be giving some Kenyan outsourced labourer reviewing the adequacy of answers a laugh :)

Re: Telling GPT-4 you're scared or under pressure improves performance

#92

> The implications are clear: incorporating emotional cues can lead to more effective and responsive AI applications. I think it's important to remember these models were trained on human interaction, and are in many ways a mirror for us to better understand human interaction. I don't think many people would be surprised if you said emotion can be used to better communicate with a human, but it is interesting to see…

This correlates with anecdata I've seen claiming that being polite with LLMs (using "please", "thank you", etc.) also improves performance. All in all, I like this. Both for what it implies about humans interacting with other humans, and for shaping norms about how best to to interact with other entities that show signs of intelligence.

Does it extend to arbitrary entities of intelligence, though, or only those whose cognitive bases are derived from human output?

Re: Telling GPT-4 you're scared or under pressure improves performance

#93

Earlier quoted context omitted.

Then why dont agents work? If those skills were real, why do they fizzle out on production data ?

That's a good question. What is different about "production data"? What do those "production people" do that suddenly makes LLMs fail on things they work well on when not "production"?

I can tell you who would love to hear your answer to that question.

Me for starters. If it works, I can quit.

Next up are Karpathy and the CTO of OpenAI. Around July and September both talked about production challenges.

AI ops was the largest subcategory of fall YC startups.

Every single ml and LLM ops individual who gets far enough deals with evals.

I don’t know, but maybe - just maybe- the issue is that people don’t understand themselves enough, to avoid assuming too much of those emergent properties.

As I recall there was also a paper that pointed out the issues with how LLMs are measured, and that the emergence of properties was not a step change once the tests were updated.

Edit- found it: https://hai.stanford.edu/news/ais-ostensible-emergent-abilit...

Re: Telling GPT-4 you're scared or under pressure improves performance

#94

I think this is the entry point needed to get peoples attention and explain: LLMs aren’t people, and emergent properties are being over extended. If LLMs are showing “better” performance when there are tokens that humans read as emotionally salient - Then the underlying text it’s trained on shows humans give better answers when emotionally salient context is provided. LLMs predict words. Any semantic validity is a si…

"A neuron just transmits signals. Any cognitive property arises as a consequence of the interplay of those electrical and chemical signals."

Do we now understand consciousness? The statement appears fundamentally limited in its implied insight.

Re: Telling GPT-4 you're scared or under pressure improves performance

#95

Am I the only one who feels bad for asking ChatGPT a "dumb" question that I know I should know, or not saying thank you when it gives me an answer? No? I'm just a weirdo? Okay. I have to push back really hard against my proclivity to humanize it, to the point where I probably don't use it as much as I should, just because I don't want to deal with the psychic stress of reminding myself that it's not a living entity.

> I have to push back really hard against this stuff Based on ChatGPT's answers, OpenAI think you should be saying thanks to it, because it helps you have a more natural conversation, which encourages you as the user to send more natural/productive prompts when asking real questions. As for asking it dumb questions, how often do you use Google as a spellcheck? I wouldn't consider this any different.

>how often do you use Google as a spellcheck? //

When Word tells me a word is wrong, but I'm almost certain it's right, and it turns out not to be in the Microsoft dictionary somehow. I mean, I have an extensive vocabulary, but not more extensive than a decent British-English dictionary.

Re: Telling GPT-4 you're scared or under pressure improves performance

#97
post #72

Earlier quoted context omitted.

> The “statistical parrot” assertion is pretty thoroughly disproven by this point errr... all NNs are just optimisations of an associative probability objective: P(Y|X), they are by definition "statistical parrots". There isn't anything to prove or disprove. People offering prompts as evidence are people who fundamentally do not understand the basics. NNs aren't strange empirical objects, they're specified by mathema…

You’re confused about what “statistical parrot” means and you don’t seem to understand the difference between an optimization objective and the resulting model. The term “parrot” is used to imply inference by something akin to a look-up table, specifically it is used to indicate poor out-of-sample performance and a lack of a proper world model. The optimization objective is irrelevant when determining the generalizat…

> it is now quite well established that GPT-4 has impressive out-of-sample performance

Err... I can show this is false, kinda trivially. People who engage in prompt-confirmation-bias aren't aware of what the in-sample is.

It's basically everything ever digitised: you can ask it for the first paragraph of every dickens novel, to what the average petal length of an iris flower is -- etc.

How are you measuring the in-sample here?

If you engage in straightfoward reasoning from first principles, and are basically aware of what the training data is, you can show in 10 seconds critical failures of generalisation.

If you want a recipe: go find some fringe api docs. Establish that it has been trained on them. Then, since they're fringe there wont be much code on github, etc. Now ask it do something non-trivial with that API. It will fail, and the mechanism will be obvious: it'll jam in correlated code that lacks relevance.

Do the same on a popular API, and see it succeed.

The in-sample will be obvious for both, and the bounday of generalisation

Re: Telling GPT-4 you're scared or under pressure improves performance

#98

Earlier quoted context omitted.

> The “statistical parrot” assertion is pretty thoroughly disproven by this point errr... all NNs are just optimisations of an associative probability objective: P(Y|X), they are by definition "statistical parrots". There isn't anything to prove or disprove. People offering prompts as evidence are people who fundamentally do not understand the basics. NNs aren't strange empirical objects, they're specified by mathema…

You’re disregarding the emergent phenomena, which are not at all understood. There was a distinct and unpredicted jump in what can loosely be described as “cognitive abilities” between GPTs 2, 3, and 4, especially after some supervised techniques like RLHF.

There is no "emergent phenomena" the pattern described is just the same as when you add +b to an ax+b model of linear data.

ie., it's just fitting capacity.

The "emergent boundary" is just an empirical measure of the necessary fitting capacity of these models on "everything ever digitised in english" given any particular functional requirement.

All the language around this area is not scientific, nor are these practices. This is superstitious neophyte engineers, hopped up on scifi, by giddy VCs who love to be told they're funding captin picard.

Re: Telling GPT-4 you're scared or under pressure improves performance

#99

Earlier quoted context omitted.

I have a whole list. Syntactic validity is not semantic validity. Word predictors not world state predictors Text prediction not fact prediction Frankly though the best answers are 1) let’s talk to infosec first 2) hey what’s the error rate ?

It's indeed a problem some people get so hyped they forget that those systems are called "language models" for a reason. They're fantastic for tasks that are as close to 100% linguistic in nature as possible, but the content might not be better than lorem ipsum in some cases, just a filler to demonstrate correct grammar. I have noticed that GPT etc had a big "wow effect" on me, the first impression can be great becau…

Yes! Super Advanced Lorem ipsum.

You can only make out when you actually push the blasted thing.

It is text gen. Just examine the premise of chain of thoughts.

Chain of thoughts promoting shouldn’t make a difference to a world model. Definitely not to a model that logic has emerged out of.

It makes a difference to a decompression function.

Edit; you may like this: https://hai.stanford.edu/news/ais-ostensible-emergent-abilit...

Re: Telling GPT-4 you're scared or under pressure improves performance

#100

Am I the only one who feels bad for asking ChatGPT a "dumb" question that I know I should know, or not saying thank you when it gives me an answer? No? I'm just a weirdo? Okay. I have to push back really hard against my proclivity to humanize it, to the point where I probably don't use it as much as I should, just because I don't want to deal with the psychic stress of reminding myself that it's not a living entity.

I have this too and to be honest, I have made the conscious decision that it is OK. I prefer to retain my habit of being polite even when it's not necessary over getting used to being rude which may then "spill" over to my human-to-human interactions.

What's interesting is that you consider short, objective focused definition of a task to be "rude".

It's like the _reported_ way in which is you say thank you to a Chinese friend they take umbrage (get angry) because it's as if you weren't expecting them to help. Whilst in other cultures not saying thank you is a big sleight.

Post reply on HN