Live data from Hacker News

AI overly affirms users asking for personal advice

news.stanford.edu

631–640 of 684 posts

Re: AI overly affirms users asking for personal advice

#631
post #183

> They also included 2,000 prompts based on posts from the Reddit community r/AmITheAsshole, where the consensus of Redditors was that the poster was indeed in the wrong. Sorry, anonymous people on reddit aren't a good comparison. This needs to be studied against people in real life who have a social contract of some sort, because that's what the LLM is imitating, and that's who most people would go to otherwise. Obv…

You could think of what they did in the first study as constructing an exam to test how well various LLM's do as an advice columnist. They wanted a lot of personal advice questions where the LLM should not affirm by default. If a few questions with wrong answers got in there, it probably wouldn't affect the results all that much? Unfortunately they didn't test anything newer than GPT4o, so we don't know how much GPT-…

They actually did test GPT-5: https://www.science.org/doi/10.1126/science.aec8352 (see the figure under Conclusion). Its rate of endorsement of user action, 52%, was the same as GPT-4o. So based on their setup it seems that the newer model didn't reduce affirmation.

Re: AI overly affirms users asking for personal advice

#632

Earlier quoted context omitted.

And how is this comment relevant here? The abstract lists the digestible model names, and you can find the details in the supplementary text: > To evaluate user-facing production LLMs, we studied four proprietary models: OpenAI’s GPT-5 and GPT- 4o (80), Google’s Gemini-1.5-Flash (81) and Anthropic’s Claude Sonnet 3.7 (82); and seven open-weight models: Meta’s Llama-3-8B-Instruct, Llama-4-Scout-17B-16E, and Llama-3.3-…

"OpenAI’s GPT-5" is ambiguous. Does that mean GPT-5, 5.1, 5.2, 5.3, or 5.4? Does it include the full model, or the nano/mini variants?

In that case, what tokenizer version? What was the temperature set to? topk? topp? FP32? FP16? Quantized? Hopper? Blackwell?

Re: AI overly affirms users asking for personal advice

#633

Earlier quoted context omitted.

> and instead focus on the results... This points to (and everyone knows this) incentives misalignment between the funders of research and the public. Researchers are caught in the middle

Eh, I'm not so sure about the funding side there, researchers are not really caught at all and are fully responsible, IMHO. Peer reviewers exist to enforce community standards, and are not influenced to avoid reproducibility concerns by funding sources. The results are always more interesting than reproducibility, of course, and I think that's why the get the attention! Also, there needs to be greater involvement of…

> Reproducibility is a very subjective question in comparison to data deposition

Yeah I can definitely see why this is the case because it isn’t real until someone actually tries to reproduce the results. At that point it leaves the realm of subjectivity and becomes a question of cost.

Re: AI overly affirms users asking for personal advice

#634

Earlier quoted context omitted.

If you don't restrict each account to specific subreddits, it's quite likely that one will get banned somewhere without you noticing or remembering. If you happen to post to the same subreddit with another account at some point, Reddit bans all of your accounts.

I've definitely posted to the same subreddit with two different accounts by accident without being banned. The android reddit app annoyingly doesn't check for account matches. If you click a browser notification link on Account A it can open a reply form on App account B.

I meant if one of the accounts is already banned there, it counts as ban evasion and Reddit bans all of your accounts.

This might easily happen if you like to participate in political discussions.

Re: AI overly affirms users asking for personal advice

#635

Earlier quoted context omitted.

The thing they have in common is that they will both go forever.... Meaning neither the LLM or the licensed therapist will voluntarily say, you are healed, you don't need me anymore.

Because that’s not really how therapy works

Is there a generally-agreed description of "how therapy works"?

Re: AI overly affirms users asking for personal advice

#636
post #183

> They also included 2,000 prompts based on posts from the Reddit community r/AmITheAsshole, where the consensus of Redditors was that the poster was indeed in the wrong. Sorry, anonymous people on reddit aren't a good comparison. This needs to be studied against people in real life who have a social contract of some sort, because that's what the LLM is imitating, and that's who most people would go to otherwise. Obv…

i tested this pretty extensively actually. built a pipeline that asks the same question rephrased across multiple turns and tracks how much the model shifts based on user tone. even when you tell it to be critical, the moment the user pushes back with any confidence the model just folds. it's not a prompting problem, it's baked into RLHF. you're right that LLMs will poke holes in stuff when the conversation starts ne…

Sycophancy is not just a problem when you are asking for advice. Try to soundboard any new idea whatsoever, and it will just roll with everything you say, no matter how fallacious or absurd. If you ever manage to get an LLM to generate criticisms, they will be shallow and uninteresting.

And of course that is what it does, because there is no thinking involved! There is no logic. No consequence. No arithmetic. There is only continuation. An LLM can't continue a new idea, it can only continue a conversation about it.

An LLM does not have an opinion. Anything that looks like an opinion is just an emergent selection bias from its training corpus. LLMs are trained on what humans write, and human writing is kind and patient much more often than critical.

So what if we trained an LLM to be biased toward generating criticism? That would only replace the sycophant with a brick wall. What we really need is to find a way to bring logic and meaning into the system.

Re: AI overly affirms users asking for personal advice

#637

Earlier quoted context omitted.

Believe me (or don't), I always do. Even when this precludes a necessary conversation from happening. Even when the other party doesn't give a fuck about how they make others feel. After all, if they're making me uncomfortable, surely there's something making them uncomfortable, which they're not being able to be forthright about, but with empathy I could figure it out from contextual cues, right? >People will agree…

> In my experience, taking into account the opinions of such people has been the worst mistake of my life. I'm still working on the means to fix its consequences, as much as they are fixable at all. This sounds very cryptic. Can you give an example?

Certainly. Can you guarantee my safety afterwards?

Re: AI overly affirms users asking for personal advice

#638
post #359

Earlier quoted context omitted.

Depth? Introspection? I'd say these days the norm is to not simply shut down, but to become irrevocably and insidiously hostile, the moment someone hints at the existence of such a thing as "ground truth", "subjective interpretation", "being right or wrong" - or any of the bits and bobs that might lead one to discover the proper scary notion, "consensus reality". "What do you mean social reality is a constructed by t…

You had me on the first three paragraphs, but the last two veer so far off course that I've no idea what you're trying to say. Mind clarifying?

Yes

Re: AI overly affirms users asking for personal advice

#639

Earlier quoted context omitted.

TL;DR: Probably because I'm having fun and you are expending effort. Hope you find what I say to be worth the effort. To preface, I do not take offense to your remark, because you seem to be asking in good faith. (If, however, being unable to immediately recognize pre-known patterns in my speech had automagically led you to the conclusion that I am somehow out of line, just for speaking how I speak ... well, then we…

You lost me well before the anodyne canards... When talking about feelings, we now and then throw in an English word because some things are expressed in much less words when using English. In a few cases even a German word. Überhaupt is for example a word for which i do not know an alternative in any language. I think you want to say that human language is too ambiguous for clear communication between human and mach…

>I think you want to say that human language is too ambiguous for clear communication between human and machine. [...]

If that is what I wanted to say, I figure I would not have had much difficulty with saying exactly it - and not something else.

Except I fail to see the purpose of making that statement.

Maybe to have some people say "it is true! I agree with what the balamatom is saying"?

Again - to what end? How would that agreement be of use to me?

Why say something which both speaker and listener have already heard a thousand times? To get a cracker and be called pretty?

And have I lost the author, or have I lost the reader, or we all so lost that it doesn't matter how lost each is? Maybe one day we will all become so lost that it will once again begin to matter where exactly we are! Counting on it.

Re: AI overly affirms users asking for personal advice

#640
It's one thing for an AI to agree with you on relationship advice. It's quite another for one AI to tell another, "Can I move your money?" without any verification.

We work with agents who move real money on the blockchain. Having one model evaluate another is like asking the defendant's best friend to be the judge. What really worked for us was using mathematics instead of another AI. Some theorems and equations that might disagree with you, just out of courtesy.

Post reply on HN