A pastime I have with papers like this is to look for the part in the paper where they say which models they tested. Very often, you find either A) it's a model from one or more years ago, only just being published now, or B) they don't even say which model they are using. Best I could find in this paper: > We evaluated 11 user-facing production LLMs: four proprietary models from OpenAI, Anthropic, and Google; and se…
And how is this comment relevant here? The abstract lists the digestible model names, and you can find the details in the supplementary text: > To evaluate user-facing production LLMs, we studied four proprietary models: OpenAI’s GPT-5 and GPT- 4o (80), Google’s Gemini-1.5-Flash (81) and Anthropic’s Claude Sonnet 3.7 (82); and seven open-weight models: Meta’s Llama-3-8B-Instruct, Llama-4-Scout-17B-16E, and Llama-3.3-…
AI overly affirms users asking for personal advice
321–330 of 684 posts
Re: AI overly affirms users asking for personal advice
#322Earlier quoted context omitted.
Pretty sure the average Redditor is AI now.
How the hell is a study on stanford.edu assuming posts on Reddit are genuine? That should be enough to get you kicked out of Stanford.
Re: AI overly affirms users asking for personal advice
#323Earlier quoted context omitted.
When you start hearing things like “you do you” or “if you know you know” it means that you went way too far. That’s a sign of discomfort. If you make uncomfortable, you won’t get diverging perspectives. People will agree to anything to get out of a social situation that makes them uncomfortable. If your goal is meaningful conversation, you may want to consider how you make people feel.
Believe me (or don't), I always do. Even when this precludes a necessary conversation from happening. Even when the other party doesn't give a fuck about how they make others feel. After all, if they're making me uncomfortable, surely there's something making them uncomfortable, which they're not being able to be forthright about, but with empathy I could figure it out from contextual cues, right? >People will agree…
This sounds very cryptic. Can you give an example?
Re: AI overly affirms users asking for personal advice
#324Even as someone who (wrongly) believed that I had high emotional intelligence, I too was bit by this. Almost a year ago when LLMs were starting to become more ubiquitous and powerful I discussed a big life/professional decision with an LLM over the course of many months. I took its recommendation. Ultimately it turned out to be the wrong decision. Thankfully it was recoverable, but it really sobered me up on LLMs. Th…
I think that if you go to an AI for advice and emotional support, it will do what most people will do - tell you what it thinks you want to hear. I am not surprised about this at all, and I do notice that when you veer into these areas, it can do it in a surprisingly subtle and dangerous way. I try to focus on results. Things like an app that does what you want, data and reports that you need, or technical things lik…
I had to deal with a close family friend going through alcohol withdrawal and getting checked in at a recovery clinic for detox and used Claude heavily. The first thing I had it do as do that “deep research” around the topic of alcohol addiction, withdrawal, etc… and then made that a project document along with clear guidelines about how it shouldn’t make inferences beyond what it in its context and supporting docs. We also spent a whole session crafting a good set of instructions (making sure it was using Anthropics own guidelines for its model…)
Little differences in prompts make a huge deal in the output.
I dunno. It is possible to use these models for dumping crazy shit you are going through. But don’t kid yourself about their output and aggressively find ways to stomp out things it has no real way to authoritatively say.
Re: AI overly affirms users asking for personal advice
#325Re: AI overly affirms users asking for personal advice
#326A pastime I have with papers like this is to look for the part in the paper where they say which models they tested. Very often, you find either A) it's a model from one or more years ago, only just being published now, or B) they don't even say which model they are using. Best I could find in this paper: > We evaluated 11 user-facing production LLMs: four proprietary models from OpenAI, Anthropic, and Google; and se…
Generally, published papers don't give a damn about reproducibility. I've seen it identified as a crisis by many. Publishers, reviewers, and researchers mostly don't care about that level of basic rigor. There's no professional repercussions or embarrassment. Agreed - if I was a reviewer for LLM papers it would be an instant rejection not listing the versions and prompts used.
While this is sadly true, it's especially true when talking about things that are stochastic in nature.
LLMs outputs, for example, are notoriously unreproducible.
Re: AI overly affirms users asking for personal advice
#327Earlier quoted context omitted.
In my local(?) community (like in my city, not my industry) there is a saying "if you had to ask for relationship advice, then you probably should break up". There is some rationale to that. People tend to hold onto relationships that don't lead anywhere in fear of "losing" what they "already have". It's probably a comfort zone thing. So if one is desperate enough to ask random strangers online about a relationship,…
> So if one is desperate enough to ask random strangers online about a relationship I'd me more inclined to ask random strangers on the internet than close friends... That said, when me and my SO had a difficult time we went to a professional. For us it helped a lot. Though as the counselor said, we were one of the few couples which came early enough. Usually she saw couples well past the point of no return. So yeah,…
Re: AI overly affirms users asking for personal advice
#328Earlier quoted context omitted.
My experience is that it tries to look at your situation in an objective way, and tries to help you to analyse your thoughts and actions. It comes across as very empathetic though, so there can lie a danger if you are easily persuaded into seeing it as a friend.
It doesn't try to do anything. It doesn't work like that. It regurgitates the most likely tokens found in the training set.
Re: AI overly affirms users asking for personal advice
#329Re: AI overly affirms users asking for personal advice
#330Earlier quoted context omitted.
This is the correct take. The advice preceded the LLM boom. They were trained on the 'dump them' advice and proceeded to reinforce the take. So why did the relationship advice change dramatically? I speculate attribution to the disinformation campaigns during this time. They were and still are grossly underestimated.
Not sure what sorts of disinformation campaigns you're referring to... There is something more interesting to consider however; the graph starts to go up in 2013, less than 6 months after the release of Tinder.