Live data from Hacker News

Investigating how prompt politeness affects LLM accuracy (2025)

arxiv.org

191–200 of 223 posts

Re: Investigating how prompt politeness affects LLM accuracy (2025)

#191

I saw this paper the other day - I feel its result may be because the "polite" prompts they have chosen arent very good at putting the ai in the roleplay-space of a valued colleague, more like a sommelier or a high-end shopkeeper. It disagrees with most other literature on the same topic, which is worth keeping in mind. This one studies gpt4o, an old model now, but a lot of other studies are on even earlier models. "…

"Can you kindly consider the following problem" seems like the most respectful of all your examples, TBH. The others sound like ass-kissing, or even sarcastic/patronizing.

well yeah - the idea is to have the model roleplay as someone who is good at solving your category of problem. You kinda have to do that a little heavy handedly unless you wanna spend all day hinting in each of your prompts

Re: Investigating how prompt politeness affects LLM accuracy (2025)

#193
post #75

Most of the comments here seem to be from people who haven’t even read the abstract, let alone the paper. The main result, mentioned in the abstract, is the opposite of what I would have guessed: > Contrary to expectations, impolite prompts consistently outperformed polite ones, with accuracy ranging from 80.8% for Very Polite prompts to 84.8% for Very Rude prompts. These findings differ from earlier studies that ass…

Even if the rude prompts are more effective, I just can't get myself to be rude in this context. Maybe it's weird but I'd rather give up that 4% accuracy increase than roleplay a dickhead

Finally, being an actual dickhead gives me that 4% edge over polite knuckleheads!

Re: Investigating how prompt politeness affects LLM accuracy (2025)

#194
post #75

Earlier quoted context omitted.

Even if the rude prompts are more effective, I just can't get myself to be rude in this context. Maybe it's weird but I'd rather give up that 4% accuracy increase than roleplay a dickhead

I do think it's odd tbh. I have some agents that return much better results with prompts like, "I'll kill your entire family if you don't return an accurate response". It's just a machine, if certain negative token inputs provide +3-10% better accuracy then I am confused why anyone would choose not to do it?

Because it tastes bad in my mouth. If I could get a 4% productivity boost by drinking a redbull, I would still choose not to drink a redbull.

Re: Investigating how prompt politeness affects LLM accuracy (2025)

#195

Earlier quoted context omitted.

> A computer is incapable of receiving it Citation please. Without examining the corpus, it's entirely possible that the training corpus has better results when you are kind to it, so one can imagine a situation where "reception of kindness" is meaningful, and in principle if you were an AI provider, you could RLHF your way to "being rude gets you worse results" as a means to train the human users.

dude... you are commenting on a research post showing a 4.8% DECLINE when being polite vs rude in prompts.

80% vs 84.8%. run a chi squares test on that

Re: Investigating how prompt politeness affects LLM accuracy (2025)

#196
post #75

Most of the comments here seem to be from people who haven’t even read the abstract, let alone the paper. The main result, mentioned in the abstract, is the opposite of what I would have guessed: > Contrary to expectations, impolite prompts consistently outperformed polite ones, with accuracy ranging from 80.8% for Very Polite prompts to 84.8% for Very Rude prompts. These findings differ from earlier studies that ass…

Even if the rude prompts are more effective, I just can't get myself to be rude in this context. Maybe it's weird but I'd rather give up that 4% accuracy increase than roleplay a dickhead

sounds like you need an AsshoLLM to sit between you and Claude to translate.

Re: Investigating how prompt politeness affects LLM accuracy (2025)

#197
post #75

Earlier quoted context omitted.

Even if the rude prompts are more effective, I just can't get myself to be rude in this context. Maybe it's weird but I'd rather give up that 4% accuracy increase than roleplay a dickhead

Ah, see, the mistake is thinking that other people are role playing…. I think rather this is how they would talk to others if they think there will be no consequences. But what do I know.

I think there's a broad spectrum of people, some of whom are role playing, some who think there are no consequences, some who have strong distinctions between the animate and inanimate, and some who just do what they think makes sense

Re: Investigating how prompt politeness affects LLM accuracy (2025)

#199

Earlier quoted context omitted.

No… the way I said it was actually deliberately obnoxious— the appropriate direct workplace response would be: “that seems oversimplified. I disagree. Here’s why:” Calling you self-absorbed added nothing of substance to the comment. It was an assumption about your mental state and a judgement of your intent based on that. There was no factual analysis or actionable insight. It was just one person explicitly stating t…

> Your assumption is reductive and self-absorbed. Bullshit. You never insulted me personally. You used strong words to disagree with my assumption, which is an important difference. It's not an insult and was not obnoxious. But I can fully understand why a person coming from an indirect culture where any criticism is taken personally would be offended and call HR overlords to punish the person giving honest opinions.…

Not considering ‘self-absorbed’ an insult reveals everything this conversation could possibly yield.

Re: Investigating how prompt politeness affects LLM accuracy (2025)

#200
post #181

Earlier quoted context omitted.

Statistical significance is about whether an effect can reliably be said to have been measured at all; it's not about whether or not the effect itself would be significant in the sense of moving some other needle. The ~5% improvement reported here might just be an artefact of the data collection or random variation, rather than a consistent repeatable change.

I know what significance means, and I also know that getting it from a p-value is nonsensical. > The ~5% improvement reported here might just be an artefact of the data collection or random variation, rather than a consistent repeatable change. You're questioning method or data representativeness, not significance. 250 samples is just about enough to for a 5% difference in NHST (stddev is around .4, so 1.64 sigma is…

Yes, it looks just barely significant. Results that are on the edge like that often aren't reproducible.
Post reply on HN