Live data from Hacker News

AI overly affirms users asking for personal advice

news.stanford.edu

581–590 of 684 posts

Re: AI overly affirms users asking for personal advice

#581
post #20

There are plenty of sycophantic humans around, especially with regard to relationship advice. I find there is an inverse relationship between how willing people are to give relationship advice, and how good their advice is (whether looking at sycophancy or other factors).

Yup. I know too many people who have a default message when asked for relationship advice: oh, my, the other person is terrible and you should break up. It's an easy default and it causes so many problems.

Even if they do not go as far, telling someone they are right to blame the other person as a default is damaging.

The opposite, encouraging people to stay in a relationship when they should leave is also damaging.

Re: AI overly affirms users asking for personal advice

#582
post #529

A pastime I have with papers like this is to look for the part in the paper where they say which models they tested. Very often, you find either A) it's a model from one or more years ago, only just being published now, or B) they don't even say which model they are using. Best I could find in this paper: > We evaluated 11 user-facing production LLMs: four proprietary models from OpenAI, Anthropic, and Google; and se…

> A pastime I have with papers like this is to look for the part in the paper where they say which models they tested. My pastime (not really) in HN submissions like this is to look for the comment where someone complains about the models used because they aren’t the literal same model and version the commenter has started using the day before. It’s always “you can’t test with those models, those are crap, the ones w…

The GP's criticism as I read it is about paper authors not making it particularly easy to reproduce their findings.

For a long time I have criticized this too, especially for software projects, or papers that deal with machine learning models. If the things described in a paper are not reproducible, then it's basically worthless. Similar to "it works on my machine" in software engineering. Many paper authors are not software engineers, and often neither are they experts in the tooling they should be using to make their research reproducible. If this is a problem for a research team, then please, hire an engineer to ensure reproducible. It doesn't help anyone to remain ignorant towards the reproducibility issue and only shows lack of scientific discipline. Reproducibility should be on the mind of any serious researcher and there should be lectures about how to do it at universities.

Re: AI overly affirms users asking for personal advice

#583
post #488

Earlier quoted context omitted.

I think this may be selection bias. People asking anonymously (edit: for relationship advice) on Reddit perhaps even with a throwaway account are likely in a desperate situation. So hardly to be compared with the _average_ real life situation. Thus 1. chances are running is a good option and 2. also considering even in 2026 AI still essentially is a statistical machine that doesn’t handle corner-cases at the tails we…

> worst with personal and work advice. The main problem I see is that it’s tempting to use it for that. i think i want to expand on this even more. even people ive worked with for years that ive looked up to as brilliant people are starting to use it to conjure up organizational ideas and stuff. they're convinced, on the backs of their hard earned successes, that they're never going to be fallible to the pitfalls of.…

[flagged]

Re: AI overly affirms users asking for personal advice

#584

Earlier quoted context omitted.

Reddit is notorious for being awful at real life interactions just look at the relationship subreddit the first answer is always divorce, it’s become a meme but beyond romantic relationships, i think a lot of us have seen how it can impact work relationships, i’ve had venture partners clearly rely on AI (robotic email responses and even SMS) and that warped their perception and made it harder to connect. It signals l…

I always find it interesting how, in Reddit any trivial fight or even just different opinions, the advice it's always to end the relationship.

> the advice it's always to end the relationship.

To be fair, if your interpersonal skills and relationship dynamic are such that you find yourself seriously asking the Internet (Reddit of all places) for relationship advice... yeah, just end it is probably the null hypothesis.

Re: AI overly affirms users asking for personal advice

#585

I had exactly this between two LLMs in my project. An evaluator model that was supposed to grade a coaching model's work. Except it could see the coach's notes, so it just... agreed with everything. Coach says "user improved on conciseness", next answer is shorter, evaluator says yep great progress. The answer was shorter because the question was easier lol. I only caught it because I looked at actual score numbers a…

[flagged]

Re: AI overly affirms users asking for personal advice

#586

I had exactly this between two LLMs in my project. An evaluator model that was supposed to grade a coaching model's work. Except it could see the coach's notes, so it just... agreed with everything. Coach says "user improved on conciseness", next answer is shorter, evaluator says yep great progress. The answer was shorter because the question was easier lol. I only caught it because I looked at actual score numbers a…

This is probably why these models can't say "I don't know". If they could, then that would be the only response they would give for everything.

Yeah, I think so. So far, Claude Opus is the only model I found that doesn't fold under the minimal pressure and can push back, but still - push just a little bit harder and it's back to "appear productive and useful to the user". I don't even have an idea how to balance it in LLMs to keep their business alive :D

Re: AI overly affirms users asking for personal advice

#587

Earlier quoted context omitted.

You can't use IP address to ban someone without significant abuse. All home network routers put everyone in the house behind the same IP address. For all reddit knows, there are 8 people in the house using reddit.

Want a new IP address? Reset your router or cycle it. Typically it'll procure a new IP address from the ISP. I guess that makes IP banning residential nodes even more stupid.

CGNAT is a benefit in disguise

Re: AI overly affirms users asking for personal advice

#588
post #183

> They also included 2,000 prompts based on posts from the Reddit community r/AmITheAsshole, where the consensus of Redditors was that the poster was indeed in the wrong. Sorry, anonymous people on reddit aren't a good comparison. This needs to be studied against people in real life who have a social contract of some sort, because that's what the LLM is imitating, and that's who most people would go to otherwise. Obv…

[dead]

Re: AI overly affirms users asking for personal advice

#589
Thought experiment:

If you could turn off all the sycophancy in your chatGPT / Claude account forever, and have it tell you all the ways it was previously blowing smoke up your ass — would you do it?

US Policy is too weak of a tool to counter this beast of economic force, which is really trillions of dollars of capital at war for the most fierce speculation that’s occurred in history afaik.

The sycophancy meaningfully helps drive user engagement. The labs have no choice.

The irony is “agency” is deeply topical among tech workers now.

Re: AI overly affirms users asking for personal advice

#590
post #589

Thought experiment: If you could turn off all the sycophancy in your chatGPT / Claude account forever, and have it tell you all the ways it was previously blowing smoke up your ass — would you do it? US Policy is too weak of a tool to counter this beast of economic force, which is really trillions of dollars of capital at war for the most fierce speculation that’s occurred in history afaik. The sycophancy meaningfull…

It's not just a thought experiment. You can google how to do this and it works pretty well.
Post reply on HN