Live data from Hacker News

Making AI chatbots friendly leads to mistakes and support of conspiracy theories

theguardian.com

11–20 of 83 posts

Re: Making AI chatbots friendly leads to mistakes and support of conspiracy theories

#11
post #8

A few weeks ago I was gently admonished by a coding agent that the code already did what I was asking it to make the code do. I was pleasantly surprised.

Betting it was Claude. That's the only LLM that will stand up to me!

Re: Making AI chatbots friendly leads to mistakes and support of conspiracy theories

#12
post #4

> “The push to make these language models behave in a more friendly manner leads to a reduction in their ability to tell hard truths and especially to push back when users have wrong ideas of what the truth might be,” said Lujain Ibrahim at the Oxford Internet Institute, the first author on the study. People aren't much different. When society pressures people to be "more friendly", eg. "less toxic" they lose their a…

> People aren't much different.

If I had a nickel for every time someone on HN responded to a criticism of LLMs with a vapid and fallacious whataboutist variation of "humans do that too!", I could fund my own AI lab.

> Why does this surprise us?

No one said they were surprised.

Re: Making AI chatbots friendly leads to mistakes and support of conspiracy theories

#13
post #8

A few weeks ago I was gently admonished by a coding agent that the code already did what I was asking it to make the code do. I was pleasantly surprised.

Betting it was Claude. That's the only LLM that will stand up to me!

In fact it was Gemini, but I don't remember which version and there are big differences. I'm signed up for all the betas and I switch among them frequently.

Re: Making AI chatbots friendly leads to mistakes and support of conspiracy theories

#14
post #4

> “The push to make these language models behave in a more friendly manner leads to a reduction in their ability to tell hard truths and especially to push back when users have wrong ideas of what the truth might be,” said Lujain Ibrahim at the Oxford Internet Institute, the first author on the study. People aren't much different. When society pressures people to be "more friendly", eg. "less toxic" they lose their a…

So Elon Musk was right in his view that Grok should focus on truth above all, even if it became offensive?

Seems like it! I find myself rather agreeing with the sentiment. The world is a offensive place, it's not gonna become less offensive from lying about it, better to stick with honesty then.

Re: Making AI chatbots friendly leads to mistakes and support of conspiracy theories

#15
post #4

> “The push to make these language models behave in a more friendly manner leads to a reduction in their ability to tell hard truths and especially to push back when users have wrong ideas of what the truth might be,” said Lujain Ibrahim at the Oxford Internet Institute, the first author on the study. People aren't much different. When society pressures people to be "more friendly", eg. "less toxic" they lose their a…

So Elon Musk was right in his view that Grok should focus on truth above all, even if it became offensive?

Grok is one of the more biased models out there.

Less truth, and more guardrails to protect musks feelings.

“Kill the boer” mean anything to you?

Re: Making AI chatbots friendly leads to mistakes and support of conspiracy theories

#17
post #4

> “The push to make these language models behave in a more friendly manner leads to a reduction in their ability to tell hard truths and especially to push back when users have wrong ideas of what the truth might be,” said Lujain Ibrahim at the Oxford Internet Institute, the first author on the study. People aren't much different. When society pressures people to be "more friendly", eg. "less toxic" they lose their a…

So Elon Musk was right in his view that Grok should focus on truth above all, even if it became offensive?

Yea, Mecha-Hitler is a real bastion of truth. /S

Re: Making AI chatbots friendly leads to mistakes and support of conspiracy theories

#18
“The Encyclopedia Galactica defines a robot as a mechanical apparatus designed to do the work of a man. The marketing division of the Sirius Cybernetics Corporation defines a robot as “Your Plastic Pal Who’s Fun to Be With.” The Hitchhiker’s Guide to the Galaxy defines the marketing division of the Sirius Cybernetics Corporation as “a bunch of mindless jerks who’ll be the first against the wall when the revolution comes,” with a footnote to the effect that the editors would welcome applications from anyone interested in taking over the post of robotics correspondent. Curiously enough, an edition of the Encyclopedia Galactica that had the good fortune to fall through a time warp from a thousand years in the future defined the marketing division of the Sirius Cybernetics Corporation as “a bunch of mindless jerks who were the first against the wall when the revolution came.”

Re: Making AI chatbots friendly leads to mistakes and support of conspiracy theories

#20
The H-neuron paper[0] found something similar (if not more general): the same bits of the model responsible for hallucination also make the model a sycophant, and also make the model easier to jailbreak.

[0] https://arxiv.org/abs/2512.01797

Post reply on HN