Live data from Hacker News

Making AI chatbots friendly leads to mistakes and support of conspiracy theories

theguardian.com

31–40 of 83 posts

Re: Making AI chatbots friendly leads to mistakes and support of conspiracy theories

#31
post #15

Earlier quoted context omitted.

So Elon Musk was right in his view that Grok should focus on truth above all, even if it became offensive?

Grok is one of the more biased models out there. Less truth, and more guardrails to protect musks feelings. “Kill the boer” mean anything to you?

[flagged]

Re: Making AI chatbots friendly leads to mistakes and support of conspiracy theories

#32

Earlier quoted context omitted.

> People aren't much different Yes they are. There is absolutely zero evidence that friendlier humans are more prone to mistakes or conspiracy theories. However, even if that were true, LLMs are not humans, anthropomorphizing them is not a helpful way to think about them.

Would be better to think of it as ‘agreeableness’ and agreeable people are more likely to shift their views to agree with those they are talking to.

I would call it obedience, and it's not the same as friendliness.

The difference, in a repeated prisoner dilemma: Friendliness is cooperating on the first move, and then conditionally. Obedience is always cooperating.

Re: Making AI chatbots friendly leads to mistakes and support of conspiracy theories

#33
I really wish they'd stop trying to suck up to me--all the "that's a really insightful question!" stuff.

I'm one of those aspy people who immediately don't trust other humans who try to fluff up my ego. Don't like it from a chatbot either.

But the fact that all the chatbots do it means that most people really crave that ego reinforcement.

Re: Making AI chatbots friendly leads to mistakes and support of conspiracy theories

#34

Earlier quoted context omitted.

Gonna set my system prompt to: "You are a Dutch person. Respond with the directness stereotypical of people from the Netherlands."

I find the LLMs target their language to the audience, so instead you could say, “I am Dutch so give it to me straight.” In my usage the LLMs gives much smarter answers when I’ve been able to convince it that I am smart enough to hear them. It doesn’t take my word for it, it seems to require evidence. I have to warm it up with some exercises where I can impress the AI. The coding focused models seem to have much lowe…

I think modern LLMs can determine if you're speaking Dutch. That's a trick that probably hasn't worked since GPT 3.

Re: Making AI chatbots friendly leads to mistakes and support of conspiracy theories

#35
post #25
post #12

Earlier quoted context omitted.

> People aren't much different. If I had a nickel for every time someone on HN responded to a criticism of LLMs with a vapid and fallacious whataboutist variation of "humans do that too!", I could fund my own AI lab. > Why does this surprise us? No one said they were surprised.

In this case I think parent-poster is trying to explain a phenomenon, rather than downplay the problem.

But it’s actively unhelpful in explaining the phenomenon, as there is no justification for equivocating LLM and human behavior. It’s just confusing and misleading.

Re: Making AI chatbots friendly leads to mistakes and support of conspiracy theories

#36
post #31
post #15

Earlier quoted context omitted.

Grok is one of the more biased models out there. Less truth, and more guardrails to protect musks feelings. “Kill the boer” mean anything to you?

[flagged]

Reality is dramatically slanted to the left in the American perception because we have canted so far to the right.

Re: Making AI chatbots friendly leads to mistakes and support of conspiracy theories

#37
post #4

> “The push to make these language models behave in a more friendly manner leads to a reduction in their ability to tell hard truths and especially to push back when users have wrong ideas of what the truth might be,” said Lujain Ibrahim at the Oxford Internet Institute, the first author on the study. People aren't much different. When society pressures people to be "more friendly", eg. "less toxic" they lose their a…

Gonna set my system prompt to: "You are a Dutch person. Respond with the directness stereotypical of people from the Netherlands."

Finnish if you want to go hard mode.

Re: Making AI chatbots friendly leads to mistakes and support of conspiracy theories

#38
post #32

Earlier quoted context omitted.

Would be better to think of it as ‘agreeableness’ and agreeable people are more likely to shift their views to agree with those they are talking to.

I would call it obedience, and it's not the same as friendliness. The difference, in a repeated prisoner dilemma: Friendliness is cooperating on the first move, and then conditionally. Obedience is always cooperating.

Agreeableness is a Big Five personality trait so a lot of the formal research into personalities uses it as one of the dimensions.

Re: Making AI chatbots friendly leads to mistakes and support of conspiracy theories

#39

Earlier quoted context omitted.

I find the LLMs target their language to the audience, so instead you could say, “I am Dutch so give it to me straight.” In my usage the LLMs gives much smarter answers when I’ve been able to convince it that I am smart enough to hear them. It doesn’t take my word for it, it seems to require evidence. I have to warm it up with some exercises where I can impress the AI. The coding focused models seem to have much lowe…

I think modern LLMs can determine if you're speaking Dutch. That's a trick that probably hasn't worked since GPT 3.

Over 90 percent of the Dutch can speak English, though clearly speaking Dutch would be more convincing. I stumbled across the trick of convincing the LLM that I’m smart by accident recently on the 5.4-Codex model. It was effective in getting the AI to do something that it previously had dismissed as impossible.

Re: Making AI chatbots friendly leads to mistakes and support of conspiracy theories

#40

I really wish they'd stop trying to suck up to me--all the "that's a really insightful question!" stuff. I'm one of those aspy people who immediately don't trust other humans who try to fluff up my ego. Don't like it from a chatbot either. But the fact that all the chatbots do it means that most people really crave that ego reinforcement.

I do have to wonder what the mix is between "our data show this is how most people want to be talked to" and "these tokens lead to better responses on objective measures of correctness." That is, in the training data insightful questions are tangled with insightful answers, so if the bot basically always treats the user like a genius it gets on the track that leads to better answers.

Or yeah, it's just people being weak to flattery.

Post reply on HN