Live data from Hacker News

Making AI chatbots friendly leads to mistakes and support of conspiracy theories

theguardian.com

21–30 of 83 posts

Re: Making AI chatbots friendly leads to mistakes and support of conspiracy theories

#21
post #4

> “The push to make these language models behave in a more friendly manner leads to a reduction in their ability to tell hard truths and especially to push back when users have wrong ideas of what the truth might be,” said Lujain Ibrahim at the Oxford Internet Institute, the first author on the study. People aren't much different. When society pressures people to be "more friendly", eg. "less toxic" they lose their a…

> People aren't much different

Yes they are. There is absolutely zero evidence that friendlier humans are more prone to mistakes or conspiracy theories.

However, even if that were true, LLMs are not humans, anthropomorphizing them is not a helpful way to think about them.

Re: Making AI chatbots friendly leads to mistakes and support of conspiracy theories

#22
post #4

> “The push to make these language models behave in a more friendly manner leads to a reduction in their ability to tell hard truths and especially to push back when users have wrong ideas of what the truth might be,” said Lujain Ibrahim at the Oxford Internet Institute, the first author on the study. People aren't much different. When society pressures people to be "more friendly", eg. "less toxic" they lose their a…

Gonna set my system prompt to: "You are a Dutch person. Respond with the directness stereotypical of people from the Netherlands."

I find the LLMs target their language to the audience, so instead you could say, “I am Dutch so give it to me straight.”

In my usage the LLMs gives much smarter answers when I’ve been able to convince it that I am smart enough to hear them. It doesn’t take my word for it, it seems to require evidence. I have to warm it up with some exercises where I can impress the AI.

The coding focused models seem to have much lower agreeableness than the chat models.

Re: Making AI chatbots friendly leads to mistakes and support of conspiracy theories

#23
post #4

> “The push to make these language models behave in a more friendly manner leads to a reduction in their ability to tell hard truths and especially to push back when users have wrong ideas of what the truth might be,” said Lujain Ibrahim at the Oxford Internet Institute, the first author on the study. People aren't much different. When society pressures people to be "more friendly", eg. "less toxic" they lose their a…

> People aren't much different Yes they are. There is absolutely zero evidence that friendlier humans are more prone to mistakes or conspiracy theories. However, even if that were true, LLMs are not humans, anthropomorphizing them is not a helpful way to think about them.

Would be better to think of it as ‘agreeableness’ and agreeable people are more likely to shift their views to agree with those they are talking to.

Re: Making AI chatbots friendly leads to mistakes and support of conspiracy theories

#24

LLM technology specifically beam-searches manifolds (or latent space) of lingustics that are closely related to the original prompt (and the pre-prompting rules of the chatbot) which it then limits its reasoning inside of. Its just the basic outcome of weights being the primary function of how it generates reasonable answers. This is the core problem with LLM tech that several researchers have been trying to figure o…

What you're saying sounds pretty cool but can you give some examples? Is this what you're talking about?

https://chatgpt.com/share/69f246e5-e0e8-83ea-aa88-6d0024b915...

Re: Making AI chatbots friendly leads to mistakes and support of conspiracy theories

#25
post #12
post #4

> “The push to make these language models behave in a more friendly manner leads to a reduction in their ability to tell hard truths and especially to push back when users have wrong ideas of what the truth might be,” said Lujain Ibrahim at the Oxford Internet Institute, the first author on the study. People aren't much different. When society pressures people to be "more friendly", eg. "less toxic" they lose their a…

> People aren't much different. If I had a nickel for every time someone on HN responded to a criticism of LLMs with a vapid and fallacious whataboutist variation of "humans do that too!", I could fund my own AI lab. > Why does this surprise us? No one said they were surprised.

In this case I think parent-poster is trying to explain a phenomenon, rather than downplay the problem.

Re: Making AI chatbots friendly leads to mistakes and support of conspiracy theories

#26

Earlier quoted context omitted.

Gonna set my system prompt to: "You are a Dutch person. Respond with the directness stereotypical of people from the Netherlands."

I find the LLMs target their language to the audience, so instead you could say, “I am Dutch so give it to me straight.” In my usage the LLMs gives much smarter answers when I’ve been able to convince it that I am smart enough to hear them. It doesn’t take my word for it, it seems to require evidence. I have to warm it up with some exercises where I can impress the AI. The coding focused models seem to have much lowe…

I'm 90 percent sure the coding agents are better in that way due to be trained on stack overflow and the LKML. Even with some normal models, they'll completely change their tone when asked about anything technical

Re: Making AI chatbots friendly leads to mistakes and support of conspiracy theories

#27
post #15

Earlier quoted context omitted.

So Elon Musk was right in his view that Grok should focus on truth above all, even if it became offensive?

Grok is one of the more biased models out there. Less truth, and more guardrails to protect musks feelings. “Kill the boer” mean anything to you?

It tells the truth, as long as you redefine truth to not include anything perceived as "liberal bias" (which by extension, also makes reality itself excluded)

Re: Making AI chatbots friendly leads to mistakes and support of conspiracy theories

#28

The H-neuron paper[0] found something similar (if not more general): the same bits of the model responsible for hallucination also make the model a sycophant, and also make the model easier to jailbreak. [0] https://arxiv.org/abs/2512.01797

Doesn't surprise me. But I don't think this is caused by friendliness, but by obedience. And I think we want the agents to be obedient. And I am afraid there is a tradeoff - more obedience means more willful ignorance of common sense ethical constraints.

Re: Making AI chatbots friendly leads to mistakes and support of conspiracy theories

#29
post #8

A few weeks ago I was gently admonished by a coding agent that the code already did what I was asking it to make the code do. I was pleasantly surprised.

Betting it was Claude. That's the only LLM that will stand up to me!

"Claude" is a big program that wraps a coding agent around a specific model. It would be the specific model that "stands up to you". I post this pedantry only because it may be helpful to you to realize this for other reasons.

Re: Making AI chatbots friendly leads to mistakes and support of conspiracy theories

#30
post #15

Earlier quoted context omitted.

So Elon Musk was right in his view that Grok should focus on truth above all, even if it became offensive?

Grok is one of the more biased models out there. Less truth, and more guardrails to protect musks feelings. “Kill the boer” mean anything to you?

Not my experience. Grok seems to be perfectly willing to roast Musk for his shortcomings.

Where did you observe the bias? Can you share any example of the conversation or post by Grok?

Post reply on HN