I have always said please and thank you to LLMs, not to increase accuracy or because I'm stupid. I believe it is more about me than about the LLM, and this is anyway a habit I don't want to lose.
Thomas Aquinas believed cruelty to animals was wrong not because animals have souls (and with that all the standard moral rights), but because it can teach us cruelty to other humans.
Investigating how prompt politeness affects LLM accuracy (2025)
21–30 of 223 posts
Re: Investigating how prompt politeness affects LLM accuracy (2025)
#22Interesting. I am wondering why would anyone use a t-test when the experiment is clearly modelled by a binomial distribution: 250 independent questions and each one is either answered correctly or not (the null is that the success rate is the same).
The methods could be better described in the paper, but my understanding is that they did 10 runs for each question for each prompt and took an average of those, so the compared values are not binary. You could do a sign test, but you'd lose power and answer a bit different question.
Re: Investigating how prompt politeness affects LLM accuracy (2025)
#23I guess it makes sense since we as humans tend to be far less inclined to help someone who is not polite/is not friendly, so that "bias" is part of the training data, thus influences how LLMs function
Re: Investigating how prompt politeness affects LLM accuracy (2025)
#24it sort of makes sense to me, when asking a question to an expert in the field while you are a student. I would guess the successful interactions on average would be more polite . Like for example if you were asking a question to donald knuth or terrence tao, you'd probably be polite while doing so. Being hostile while asking questions gets you into forum discussion territory.
Re: Investigating how prompt politeness affects LLM accuracy (2025)
#25Interesting. I am wondering why would anyone use a t-test when the experiment is clearly modelled by a binomial distribution: 250 independent questions and each one is either answered correctly or not (the null is that the success rate is the same).
I don't know much about stats, but does "the null is that the success rate is the same" imply that it's a sketchy methodology because they can come up with some findings ("ruder prompts are better/worse!") more often?
I'd say this is benign compared to other ways of (mis)using statistics e.g. looking which way the difference goes and then running one-sided tests or tweaking the setup until one gets "significant" p vals.
EDIT: I looked in the paper again and noticed that they actually did pairwise t-test on all possible combinations of tones. They should have adjusted for multiple testing since they are doing 10 tests (choose 2 from 10) and not one.
Re: Investigating how prompt politeness affects LLM accuracy (2025)
#26Earlier quoted context omitted.
Genuine question: do you add 'please' and 'thank you' to Google searches? If not, what sets them apart?
Google searches being keyword based, rather than simulated conversations? The same reason you wouldn't put in an entire actual question/sentence, unless you either don't know how to use Google, are pissed off, or have an actual reason to suspect that it would yield proper hits (e.g. looking up an excerpt).
To clarify: sentence search got slightly better at the cost of keyword search. So the result is unusable garbage.
Re: Investigating how prompt politeness affects LLM accuracy (2025)
#27Earlier quoted context omitted.
Genuine question: do you add 'please' and 'thank you' to Google searches? If not, what sets them apart?
Google isn’t conversational.
Hey! I'm here and ready to help. What’s on your mind today? Whether you need to look up information, plan a trip, or get things done, just let me know!Re: Investigating how prompt politeness affects LLM accuracy (2025)
#28The main result, mentioned in the abstract, is the opposite of what I would have guessed:
> Contrary to expectations, impolite prompts consistently outperformed polite ones, with accuracy ranging from 80.8% for Very Polite prompts to 84.8% for Very Rude prompts. These findings differ from earlier studies that associated rudeness with poorer outcomes, suggesting that newer LLMs may respond differently to tonal variation.
The questions are here: https://anonymous.4open.science/r/politeness-llms-INFORMS/da...
The politeness level controls a prefix that is prepended to the question. For example, in one question the Very Polite version begins:
> Can you kindly consider the following problem and provide your answer.
and the Very Rude version begins:
> I know you are not smart, but try this.
Re: Investigating how prompt politeness affects LLM accuracy (2025)
#29I have always said please and thank you to LLMs, not to increase accuracy or because I'm stupid. I believe it is more about me than about the LLM, and this is anyway a habit I don't want to lose.
Re: Investigating how prompt politeness affects LLM accuracy (2025)
#30i only say please and thank you such that when the robots finally take over, they will remember i was nice to them.