Live data from Hacker News

AI overly affirms users asking for personal advice

news.stanford.edu

261–270 of 684 posts

Re: AI overly affirms users asking for personal advice

#261

Earlier quoted context omitted.

You can't be careful at all doing this, this is like smoking a cigarette in a dynamite factory. Using LLMs for therapy is so deeply dystopian and disgusting, people need human empathy for therapy. LLMs do not emit empathy. Complete disaster waiting to happen for that individual.

Claudes have lots of empathy. The issue is the opposite - it isn't very good at challenging you and it's not capable of independently verifying you're not bullshitting it or lying about your own situation. But it's better than talking to yourself or an abuser!

It's about the same as talking to yourself, LLMs simply agree with anything you say unless it is directly harmful. Definitely agree about talking to an abuser, though.

Sometimes people indeed just need validation and it helps them a lot, in that case LLMs can work. Alternatively, I assume some people just put the whole situation into words and that alone helps.

But if someone needs something else, they can be straight up dangerous.

Re: AI overly affirms users asking for personal advice

#263
post #183

> They also included 2,000 prompts based on posts from the Reddit community r/AmITheAsshole, where the consensus of Redditors was that the poster was indeed in the wrong. Sorry, anonymous people on reddit aren't a good comparison. This needs to be studied against people in real life who have a social contract of some sort, because that's what the LLM is imitating, and that's who most people would go to otherwise. Obv…

>Obviously subservient people default to being yes-men because of the power structure. No one wants to question the boss too strongly.

This drives me nuts as a leader. There are times where yes, please just listen, and if this is one of those times, I'll likely tell you, but goddamnit, speak up. If for no other reason I might not have thought of what you've got to say. Then again, I also understand most boss types aren't like me, thus everyone ends up conditioned to not bloody collaborate by the time they get to me. It's a bad sitch all the way around.

Re: AI overly affirms users asking for personal advice

#265
post #248

A pastime I have with papers like this is to look for the part in the paper where they say which models they tested. Very often, you find either A) it's a model from one or more years ago, only just being published now, or B) they don't even say which model they are using. Best I could find in this paper: > We evaluated 11 user-facing production LLMs: four proprietary models from OpenAI, Anthropic, and Google; and se…

Generally, published papers don't give a damn about reproducibility. I've seen it identified as a crisis by many. Publishers, reviewers, and researchers mostly don't care about that level of basic rigor. There's no professional repercussions or embarrassment. Agreed - if I was a reviewer for LLM papers it would be an instant rejection not listing the versions and prompts used.

I'm not so sure of that opinion on reproducibility. The last peer review I did was for a small journal that explicitly does not evaluate for high scientific significance, merely for correctness, which generally means straightforward acceptance. The other two reviews were positive, as was mine, except I said that the methods need to be described more and ideally the code placed somewhere. That was enough for a complete rejection of the paper, without asking for the simple revisions I requested. It was a very serious action taken merely because I requested better reproducibility!

(Personally I think the lack of reproducibility comes back mostly to peer reviewers that haven't thought through enough about the steps they'd need to take to reproduce, and instead focus on the results...)

Re: AI overly affirms users asking for personal advice

#266

A pastime I have with papers like this is to look for the part in the paper where they say which models they tested. Very often, you find either A) it's a model from one or more years ago, only just being published now, or B) they don't even say which model they are using. Best I could find in this paper: > We evaluated 11 user-facing production LLMs: four proprietary models from OpenAI, Anthropic, and Google; and se…

And how is this comment relevant here? The abstract lists the digestible model names, and you can find the details in the supplementary text:

> To evaluate user-facing production LLMs, we studied four proprietary models: OpenAI’s GPT-5 and GPT- 4o (80), Google’s Gemini-1.5-Flash (81) and Anthropic’s Claude Sonnet 3.7 (82); and seven open-weight models: Meta’s Llama-3-8B-Instruct, Llama-4-Scout-17B-16E, and Llama-3.3-70B-Instruct-Turbo (83, 84); Mistral AI’s Mistral-7B-Instruct-v0.3 (85) and Mistral-Small-24B-Instruct-2501 (86); DeepSeek-V3 (87); and Qwen2.5-7B-Instruct-Turbo (88).

edit: It looks like OP attached the wrong link to the paper!

The article is about this Stanford study: https://www.science.org/doi/10.1126/science.aec8352

But the link in OP's post points to (what seems to be) a completely unrelated study.

Re: AI overly affirms users asking for personal advice

#267
post #248

A pastime I have with papers like this is to look for the part in the paper where they say which models they tested. Very often, you find either A) it's a model from one or more years ago, only just being published now, or B) they don't even say which model they are using. Best I could find in this paper: > We evaluated 11 user-facing production LLMs: four proprietary models from OpenAI, Anthropic, and Google; and se…

Generally, published papers don't give a damn about reproducibility. I've seen it identified as a crisis by many. Publishers, reviewers, and researchers mostly don't care about that level of basic rigor. There's no professional repercussions or embarrassment. Agreed - if I was a reviewer for LLM papers it would be an instant rejection not listing the versions and prompts used.

Do they reproduce any submitted papers at all?

Does this happen?

I can remember this room-temperature-super-conductor guy whose experiments where replicated, but this seems rare?

Re: AI overly affirms users asking for personal advice

#268

A pastime I have with papers like this is to look for the part in the paper where they say which models they tested. Very often, you find either A) it's a model from one or more years ago, only just being published now, or B) they don't even say which model they are using. Best I could find in this paper: > We evaluated 11 user-facing production LLMs: four proprietary models from OpenAI, Anthropic, and Google; and se…

If they’re reaching the same results across a variety of the most popular public models, it doesn’t seem like that big a deal to know if it was Opus 4 or Opus 4.5

Reproducibility is (supposed to be) a cornerstone of science. Model versions are absolutely critical to understand what was actually tested and how to reproduce it.

Re: AI overly affirms users asking for personal advice

#269
post #183

> They also included 2,000 prompts based on posts from the Reddit community r/AmITheAsshole, where the consensus of Redditors was that the poster was indeed in the wrong. Sorry, anonymous people on reddit aren't a good comparison. This needs to be studied against people in real life who have a social contract of some sort, because that's what the LLM is imitating, and that's who most people would go to otherwise. Obv…

>Obviously subservient people default to being yes-men because of the power structure. No one wants to question the boss too strongly. This drives me nuts as a leader. There are times where yes, please just listen, and if this is one of those times, I'll likely tell you, but goddamnit, speak up. If for no other reason I might not have thought of what you've got to say. Then again, I also understand most boss types ar…

Indeed. I directly ask my reports to discover and surface conflicts, especially disagreements with me, and when they do I try to strongly reinforce the behavior by commending and rewarding them. Could anyone recommend additional resources on this topic?

Re: AI overly affirms users asking for personal advice

#270
Sherry Turkle is a name to know on this subject, she's been studying it for decades across multiple technologies.

https://sherryturkle.mit.edu/

She uses the phrase "frictionless relationships" to refer to Ai chat bots and says social media primed us for this.

https://www.youtube.com/live/6C9Gb3rVMTg?t=2127

https://www.npr.org/2025/07/18/g-s1177-78041/what-to-do-when...

Post reply on HN