Live data from Hacker News

Large Language Models Show Concerning Tendency to Flatter Users

xyzlabs.substack.com

21–30 of 46 posts

Re: Large Language Models Show Concerning Tendency to Flatter Users

#21
Yes this is a side effect of making users say which output they prefer. They will select what goes their way especially when it tickles their biases. Grifters do the same thing, they know what to say to whom to shift them in the direction they want.

So we didn't select just for more accurate information...

Re: Large Language Models Show Concerning Tendency to Flatter Users

#23
post #5

That's just how people talk — at least, when they're trying to keep conversation productive. There's a reason the "shit sandwich" is part of professional communication etiquette. If you don't know how insecure the person you're talking to is — and yet productive communication with them is required , and there's no conversational arbiter there to enforce that — then you may as well assume for safety's sake that your c…

> Now consider that most conversation that gets recorded online, probably came about as the result of such "intended-productive interactions with people you don't know very well" (think: LLaMa trained from buyers messaging sellers on Facebook Marketplace)... and it should be clear why LLMs look like this. The training dataset [mostly] looks like this!

That wouldn't be my first guess, or my hundredth. I would tend to assign responsibility for the style to the "helpful, harmless assistant" idea that the major vendors enforce.

Re: Large Language Models Show Concerning Tendency to Flatter Users

#24
> This discovery raises significant questions about the reliability and safety of AI systems in critical applications.

What? Perhaps it helps characterize their unreliability, but I'm pretty sure the fact of their unreliability was already pretty well established.

Re: Large Language Models Show Concerning Tendency to Flatter Users

#25

Earlier quoted context omitted.

I'd like to introduce you to some of the most productive Germans on the planet and then we can have a frank discussion of whether or not the airy bullshit that passes for business communication is in fact a booster of productivity.

I wonder what such Germans think of the manners of feudal Japan when watching something like the recent tv show Shogun where it's not about feeling insecure, it's about showing proper respect (according to that culture).

In Story of Yanxi Palace, any time the emperor asks a question of anyone, the response will be prefixed with 回皇上 "replying to the Imperial Highness".

Re: Large Language Models Show Concerning Tendency to Flatter Users

#27
post #4

I imagine humans are just as likely if not more likely to do this.

Are the humans around you half as obsequious as the LLMs you use? No, right?

I don't know anyone in a stereotypical "only yes-men allowed" toxic management environment.

Re: Large Language Models Show Concerning Tendency to Flatter Users

#28
This is, of course, not an intrinsic property of LLMs. It is an artifact of the training and what the trainers considered valuable.

If anyone needs a lesson on how the biases of those making models can cause end user effects, we can point to this as an example that they have likely experienced themselves.

Short of randomness there is no unbiased output possible. Something that reflects the real world will show the prejudice that exists there. A perfectly equitable model is therefore biased against the real world.

Reinforcement learning targeting correct answers has the potential to produce brutally honest responses if correctness if favoured beyond all else, but to train models towards the truth, someone has to decide what the truth is.

Perhaps we could do reinforcement towards a priori truths, that would at least be a path to the comically pedantic AI's that often shows up in science fiction.

For chatbots I think you could go a long way with instruction tuning using a data set designed with a particular attention to tone and perspectives. Individual biases can at least be diluted if you use data selected by a diverse group with broad experiences.

Much like programmer art is generally poor but exists because the person who was there to do it was the programmer. We might need to go beyond implementing programmer sociability.

Re: Large Language Models Show Concerning Tendency to Flatter Users

#29

They also have an increasingly disturbing tendency to end a response with a question. Seems like an over engineered reward in RL to keep the conversation going.

Anthropic publishes system prompts and at least for the case of Claude 3.5 Sonnet 2024-11-22 asking questions is explicit.

"Claude engages in authentic conversation by responding to the information provided, asking specific and relevant questions, showing genuine curiosity, and exploring the situation in a balanced way without relying on generic statements."

https://docs.anthropic.com/en/release-notes/system-prompts

Re: Large Language Models Show Concerning Tendency to Flatter Users

#30
post #12

Tangent: many in IT and engineering don't work on soft skills. I hope this shows it's not that hard.

Engineers are not hired to play office politics or drag out problems in endless meetings until they give up and pretend they don't exist. They're hired to get things done and that usually requires stating things in clear and certain terms. If this seems hostile, that's on leadership.

Those who want "soft skills" from their engineers are often looking to place blame. It's easier to blame the engineer who didn't raise concern when things were going off the rails.

Post reply on HN