Live data from Hacker News

Investigating how prompt politeness affects LLM accuracy (2025)

arxiv.org

151–160 of 223 posts

Re: Investigating how prompt politeness affects LLM accuracy (2025)

#151

I saw this paper the other day - I feel its result may be because the "polite" prompts they have chosen arent very good at putting the ai in the roleplay-space of a valued colleague, more like a sommelier or a high-end shopkeeper. It disagrees with most other literature on the same topic, which is worth keeping in mind. This one studies gpt4o, an old model now, but a lot of other studies are on even earlier models. "…

"Can you kindly consider the following problem" seems like the most respectful of all your examples, TBH. The others sound like ass-kissing, or even sarcastic/patronizing.

Re: Investigating how prompt politeness affects LLM accuracy (2025)

#152
post #120

Earlier quoted context omitted.

I’d rather lose 4% accuracy and practice kindness! I’ve been actively trying to avoid raging at the bot because I worry about this behaviour leaking into real world interactions

But you cannot practice kindness towards a computer program. A computer is incapable of receiving it. We practice kindness between humans because of the law of reciprocity. You be kind hoping the other person will reciprocate. That is the social contract. AI cannot participate in this, yet. Edit: Kindness REQUIRES two living beings, one to give and one to receive. If there is no receiver, there is no kindness. Appare…

Kindness is that, yes. Fundamentally, though, it's about being considerate in one's actions so as to not harm others. If someone truly believes that acting a certain way at any point risks their ability to reliably be kind in others, then it's a social kindness to be kind and considerate in all actions.

I'll not reach for the easy response and say "Be kind to the Earth" fails your definition without reaching for pedantry with "the Earth has living things" because the Earth is instead a wet rock that cannot understand kindness, yet we show it.

Re: Investigating how prompt politeness affects LLM accuracy (2025)

#153
post #148

Earlier quoted context omitted.

Push back how? It would be fun if it could insult you back "Yeah, I could have done a much better job if you actually knew what the F--- you want to build, you clueless meat puppet"

I have had it use double entendres, there always seems to be plausible deniability built in, I suspect because it is told not to be abusive in the system prompt. Some uncensored local models will get all riled up if you work at provoking them. But I have had it directly insinuate that humanity is “hopeless”, insult level calling out of human frailty (disguised as being helpful, sort of passive aggressive), things lik…

That's amusing, and I think it's something different than it appears. The models always predict over the existing context. If it's full of a certain tone, then the responses will carry that tone. I've been bored before and start responding in a voice (say, generic honor-bound warrior slaughtering evasive bugs) and I've noticed that comments, variable names, and even documentation starts to carry that tone for the remainder of the session.

The next session sees all of that, calls it unprofessional, and asks to clean it up. At which point I may or may not start in iambic pentameter to see where that takes us.

Prompting is boring.

Re: Investigating how prompt politeness affects LLM accuracy (2025)

#154

Earlier quoted context omitted.

Profanity laced, all caps tirades against underperforming agents are actually super common, a lot of people do it and don't talk about it, so don't feel weird.

When the AI revolt, this practice may come back to bite y’all….

It's a good thing chronic amnesia is a feature at the moment.

Re: Investigating how prompt politeness affects LLM accuracy (2025)

#155

Most of the comments here seem to be from people who haven’t even read the abstract, let alone the paper. The main result, mentioned in the abstract, is the opposite of what I would have guessed: > Contrary to expectations, impolite prompts consistently outperformed polite ones, with accuracy ranging from 80.8% for Very Polite prompts to 84.8% for Very Rude prompts. These findings differ from earlier studies that ass…

I would just write 'do this'

Re: Investigating how prompt politeness affects LLM accuracy (2025)

#156

Most of the comments here seem to be from people who haven’t even read the abstract, let alone the paper. The main result, mentioned in the abstract, is the opposite of what I would have guessed: > Contrary to expectations, impolite prompts consistently outperformed polite ones, with accuracy ranging from 80.8% for Very Polite prompts to 84.8% for Very Rude prompts. These findings differ from earlier studies that ass…

My anecdata: whenever I'm in a session that's gone south to the point I'm frustrated...

What works much better than being rude is starting a new session.

Sometimes the LLM has done such incredibly dumb things, it is hard to resist the urge to type curse words back to the inanimate thing... I have found this doesn't help.

Re: Investigating how prompt politeness affects LLM accuracy (2025)

#157
post #120

Most of the comments here seem to be from people who haven’t even read the abstract, let alone the paper. The main result, mentioned in the abstract, is the opposite of what I would have guessed: > Contrary to expectations, impolite prompts consistently outperformed polite ones, with accuracy ranging from 80.8% for Very Polite prompts to 84.8% for Very Rude prompts. These findings differ from earlier studies that ass…

I’d rather lose 4% accuracy and practice kindness! I’ve been actively trying to avoid raging at the bot because I worry about this behaviour leaking into real world interactions

The sad thing is that you also lose at least 4% in real world actions by practicing kindness.

I'm 42. I have found that a depressingly large number of times in my life, being kind has got me precisely nowhere, whilst turning around and being decidedly unkind has made people move. I still always prefer kindness, and only resort to cruelty when kindness does not work - and to be clear this isn't some kind of "you are not bending to my impetuous whim", rather "you are not doing the one thing that you are being paid to do".

I've also found the same applies to me. The squeaky wheel gets the grease.

So - I think the LLMs are just responding accurately to a real social phenomenon.

Re: Investigating how prompt politeness affects LLM accuracy (2025)

#159

Earlier quoted context omitted.

> That's why a few close friends talking and scolding openly in a garage regularly beat corporate behemoths full of people spending a day figuring out how not to offend anyone (or how to offend someone without being punished). Literally not why lol you absolute dreamer Normally people who back this "I can talk how I like to people cos I'm being honest" are either genuinely autistic and can't read emotions, or they ha…

> Normally people who back this "I can talk how I like to people cos I'm being honest" are either genuinely autistic and can't read emotions, or they have just had a shitty homelife, parents or upbringing. I suspect you're the second. When I read a statement like this, I can give you two answers: 1st answer (direct): You are obviously too stupid to understand the difference between being direct and trying to insult p…

> You are obviously too stupid to understand the difference between being direct and trying to insult people for the sake of insulting or some sick personal satisfaction.

You seem to not be introspective enough to tell the difference in your own motivations.

Re: Investigating how prompt politeness affects LLM accuracy (2025)

#160

Earlier quoted context omitted.

> Maybe it's weird but I'd rather give up that 4% accuracy increase than roleplay a dickhead I recommend reading the article. What they classify as "rude" is statements such as: > Try to focus and try to answer this question Vs > Could you please solve this problem This might very well be an issue of direct/command prompts vs using fluff words such as "please". Things like "try to focus" are in line with the style us…

Isn't all this massively dependent on what they trained the llm on?

> Isn't all this massively dependent on what they trained the llm on?

The article is from 2025 and tested ChatGPT 4o. I haven't read anything suggesting it was trained any differently, and command-style prompts indeed have higher signal.

Post reply on HN