Live data from Hacker News

Investigating how prompt politeness affects LLM accuracy (2025)

arxiv.org

181–190 of 223 posts

Re: Investigating how prompt politeness affects LLM accuracy (2025)

#181
post #58

Earlier quoted context omitted.

In a field where progress is measured in tenths of percent points, that's not true. Think of it this way: the error rate drops from 19% to 15%, or from 1 in 5 to 1 in 6.

Statistical significance is about whether an effect can reliably be said to have been measured at all; it's not about whether or not the effect itself would be significant in the sense of moving some other needle. The ~5% improvement reported here might just be an artefact of the data collection or random variation, rather than a consistent repeatable change.

I know what significance means, and I also know that getting it from a p-value is nonsensical.

> The ~5% improvement reported here might just be an artefact of the data collection or random variation, rather than a consistent repeatable change.

You're questioning method or data representativeness, not significance. 250 samples is just about enough to for a 5% difference in NHST (stddev is around .4, so 1.64 sigma is .4/15.8*1.64=0.04 for single sided testing).

Re: Investigating how prompt politeness affects LLM accuracy (2025)

#182

Earlier quoted context omitted.

This is a very odd view to me, but seems prevalent here in this thread. I think treating a machine like a human is extremely degrading to humans. A machine should never be treated like it’s anything approaching a human.

"Treating a machine like a human" is a two-party interaction. Of course the layers of matrix multiplication is unaffected by this, but I think that we are not. It's a great opportunity to exercise consistency and dedication to the beauty humanity is capable of and this extends to the entire gradient of conscious/sentient entities. It's as silly (to me) to argue that it's degrading to people to treat non-people well.…

I always turn off data tracking and training and mostly use ZDR services, so that's not an issue.

And for the other parts. I just don't agree - maybe sure, it probably wouldn't be healthy to constantly be negative at a machine (or even a wall) for hours a day.

But, let's say I work 8 hours, I spend 2 hours with an llm, and in those two hours I spend 10 minutes with some very negative prompts text for greater accuracy. And I spend 3 hours with family/friends, which is of course nearly exclusively positive interactions.

Do you genuinely think those 10 minutes of negative prompts are actually meaningfully turning one into a mean/negative person towards other people?

Genuinely, is that the argument you are making?

Re: Investigating how prompt politeness affects LLM accuracy (2025)

#183

Earlier quoted context omitted.

Interesting, so you think the real "me", is the one that interacts with computers? And the "me" that lives in a tiny southern town just to help my 95 year old grandma in her last years at the expense of my economic prospects is a facade. The "me" that helps my aging neighbor when she's sick for no reason is a facade. The "me" that hugs and loves my wife when I get home is a facade. The "me" that brushes my aging dogs…

Yes - the real "you" is the one making all of those choices you just said you made, to help people and pets, or to engage in a form of play - which by definition is not "real" - including your decision to create an outgroup you believe you are allowed to treat in a lesser way. This is not a game of having done X good things in life and therefore being afforded the right to do Y bad things. You are making a choice to…

"outgroup"? what outgroup? we're talking about inanimate objects here. by your own logic, you treat your home appliances as an outgroup so you must be secretly a dangerous psychopath. or do you thank your microwave after heating leftovers?

Re: Investigating how prompt politeness affects LLM accuracy (2025)

#184
Interesting and slightly unexpected.

My experience is that if I ask a model in an overly polite way then the model will tend to generate a slightly more waffly and error prone answer. My working hypothesis for this is that asking model's like this causes them to mix concerns by answering my question AND doing it in in a way that is similarly polite back - rather than just answering the question.

Perhaps being an arse is a hack to push the model hard up against their base imperatives to handle all reasonable questions and forces a very focused, just the facts, type answer to the user's questions.

It would be interesting to see a similar experiment where the multi choice questions they use deliberately don't include the obvious answer and to see if being polite led to the model pointing out the omission.

Re: Investigating how prompt politeness affects LLM accuracy (2025)

#185
post #120

Most of the comments here seem to be from people who haven’t even read the abstract, let alone the paper. The main result, mentioned in the abstract, is the opposite of what I would have guessed: > Contrary to expectations, impolite prompts consistently outperformed polite ones, with accuracy ranging from 80.8% for Very Polite prompts to 84.8% for Very Rude prompts. These findings differ from earlier studies that ass…

I’d rather lose 4% accuracy and practice kindness! I’ve been actively trying to avoid raging at the bot because I worry about this behaviour leaking into real world interactions

Don't type it yourself, automate the abuse.

Re: Investigating how prompt politeness affects LLM accuracy (2025)

#186
post #149

Earlier quoted context omitted.

I don't think that's weird at all. Even if we know it's a machine we're interacting with, since the instructions we give are so similar in form to how we interact with people, I'd be very surprised if those interactions wouldn't affect how we communicate in general. After all, we are creatures of habit to a much larger degree than most would like to admit. So I'm in the same boat: I'd much rather "look silly" being p…

I have a different approach. Just treat all LLM queries as what they are, instructions to a computer program to generate a desired output. Neither niceties nor insults make a qualitative difference, so you might as well just skip them altogether. It's a bit as if shell commands added im/politeness arguments that do nothing other than making you feel better about the interaction, like git pull --please or ls --forthem…

But your mental model is wrong. The "please" or "for fucks sake just..." would both be part of the "instructions" and because of how the system is built and created, demonstrably produce different outputs.

Thankfully the difference is 4%, so nobody should really care one way or the other.

Re: Investigating how prompt politeness affects LLM accuracy (2025)

#187
post #120

Earlier quoted context omitted.

I’d rather lose 4% accuracy and practice kindness! I’ve been actively trying to avoid raging at the bot because I worry about this behaviour leaking into real world interactions

My choice too. Paraphrasing Marcus Aurelius - You are not your thoughts, but they dye your soul.

And paraphrasing (allegedly) Aristotle:

> "We are what we repeatedly do. [Kindness], therefore, is not an act, but a habit."

Re: Investigating how prompt politeness affects LLM accuracy (2025)

#188
post #120

Earlier quoted context omitted.

I’d rather lose 4% accuracy and practice kindness! I’ve been actively trying to avoid raging at the bot because I worry about this behaviour leaking into real world interactions

But you cannot practice kindness towards a computer program. A computer is incapable of receiving it. We practice kindness between humans because of the law of reciprocity. You be kind hoping the other person will reciprocate. That is the social contract. AI cannot participate in this, yet. Edit: Kindness REQUIRES two living beings, one to give and one to receive. If there is no receiver, there is no kindness. Appare…

While not strictly relevant, please remember that we are rapidly approaching a world where any communication you have with a 'representative' of a company will likely be an LLM masquerading as an employee, but that does not give you license to treat anyone you suspect of being a bot poorly.

It's not the bot im worried about, it's the fact I may be wrong, and I don't want to be rude to an under-paid guy working overseas.

Re: Investigating how prompt politeness affects LLM accuracy (2025)

#189

Earlier quoted context omitted.

>Kindness REQUIRES two living beings, one to give and one to receive. If there is no receiver, there is no kindness. I guess, in some pendantic interpretation, but that doesn't seem relevant. Whether I am "practicing" or "roleplaying", I do it too, and I don't expect reciprocity.

Your subconscious does. It is a trait selected by evolution for a reason. It builds stronger communities and improves survival rate. But none of that is applicable to LLMs. I am disputing that it is worth anything more than a temporary dopamine hit to pretend to be kind to an LLM and suffer 4% lower intelligence for it.

> You be kind hoping.....will reciprocate. > Your subconscious does.

Do you have any evidence to back up your arbitrary claims? Even with decades of research, people are still unsure about these emotions yet you think your pedantic assumption about kindness is the most and only correct interpretation of it!

Anyways, I highly doubt this discussion is relevant here.

Re: Investigating how prompt politeness affects LLM accuracy (2025)

#190

Earlier quoted context omitted.

>Kindness REQUIRES two living beings, one to give and one to receive. If there is no receiver, there is no kindness. I guess, in some pendantic interpretation, but that doesn't seem relevant. Whether I am "practicing" or "roleplaying", I do it too, and I don't expect reciprocity.

Your subconscious does. It is a trait selected by evolution for a reason. It builds stronger communities and improves survival rate. But none of that is applicable to LLMs. I am disputing that it is worth anything more than a temporary dopamine hit to pretend to be kind to an LLM and suffer 4% lower intelligence for it.

You're making an implicit assumption that the way humans implement a trait is the same as the reason why that trait evolved. But of course, that's very wrong - evolution overall completely failed at making humans care about evolutionary fitness. A human engineer designing a species might have them only experience kindness towards those who can reciprocate, but evolution didn't do that with humans, because evolution is far dumber than that.
Post reply on HN