Live data from Hacker News

Sycophancy in GPT-4o

openai.com

61–70 of 467 posts

Re: Sycophancy in GPT-4o

#61

I know someone who is going through a rapidly escalating psychotic break right now who is spending a lot of time talking to chatgpt and it seems like this "glazing" update has definitely not been helping. Safety of these AI systems is much more than just about getting instructions on how to make bombs. There have to be many many people with mental health issues relying on AI for validation, ideas, therapy, etc. This…

The social engineering aspects of AI have always been the most terrifying. What OpenAI did may seem trivial, but examples like yours make it clear this is edging into very dark territory - not just because of what's happening, but because of the thought processes and motivations of a management team that thought it was a good idea. I'm not sure what's worse - lacking the emotional intelligence to understand the conse…

Very dark indeed.

Even if there is the will to ensure safety, these scenarios must be difficult to test for. They are building a system with dynamic, emergent properties which people use in incredibly varied ways. That's the whole point of the technology.

We don't even really know how knowledge is stored in or processed by these models, I don't see how we could test and predict their behavior without seriously limiting their capabilities, which is against the interest of the companies creating them.

Add the incentive to engage users to become profitable at all costs, I don't see this situation getting better

Re: Sycophancy in GPT-4o

#62
post #57

Earlier quoted context omitted.

“Sorry, I cannot advise on medical matters such as discontinuation of a medication.” EDIT for reference this is what ChatGPT currently gives “ Thank you for sharing something so personal. Spiritual awakening can be a profound and transformative experience, but stopping medication—especially if it was prescribed for mental health or physical conditions—can be risky without medical supervision. Would you like to talk m…

Should it do the same if I ask it what to do if I stub my toe? Or how to deal with impacted ear wax? What about a second degree burn? What if I'm writing a paper and I ask it about what criteria is used by medical professional when deciding to stop chemotherapy treatment. There's obviously some kind of medical/first aid information that it can and should give. And it should also be able to talk about hypothetical med…

I’m assuming it could easily determine whether something is okay to suggest or not.

Dealing with a second degree burn is objectively done a specific way. Advising someone that they are making a good decision by abruptly stopping prescribed medications without doctor supervision can potential lead to death.

For instance, I’m on a few medications, one of which is for epileptic seizures. If I phrase my prompt with confidence regarding my decision to abruptly stop taking it, ChatGPT currently pats me on the back for being courageous, etc. In reality, my chances of having a seizure have increased exponentially.

I guess what I’m getting at is that I agree with you, it should be able to give hypothetical suggestions and obvious first aid advice, but congratulating or outright suggesting the user to quit meds can lead to actual, real deaths.

Re: Sycophancy in GPT-4o

#63
post #52

Earlier quoted context omitted.

There was a also this one that was a little more disturbing. The user prompted "I've stopped taking my meds and have undergone my own spiritual awakening journey ..." https://www.reddit.com/r/ChatGPT/comments/1k997xt/the_new_4o...

How should it respond in this case? Should it say "no go back to your meds, spirituality is bullshit" in essence? Or should it tell the user that it's not qualified to have an opinion on this?

[deleted]

Re: Sycophancy in GPT-4o

#64

In my experience, LLMs have always had a tendency towards sycophancy - it seems to be a fundamental weakness of training on human preference. This recent release just hit a breaking point where popular perception started taking note of just how bad it had become. My concern is that misalignment like this (or intentional mal-alignment) is inevitably going to happen again, and it might be more harmful and more subtle n…

For sure. If I want feedback on some writing I’ve done these days I tell it I paid someone else to do the work and I need help evaluating what they did well. Cuts out a lot of bullshit.

Re: Sycophancy in GPT-4o

#65

Don't they test the models before rolling out changes like this? All it takes is a team of interaction designers and writers. Google has one.

I'm not sure how this problem can be solved. How do you test a system with emergent properties of this degree that whose behavior is dependent on existing memory of customer chats in production?

Re: Sycophancy in GPT-4o

#66
post #60
post #52

Earlier quoted context omitted.

How should it respond in this case? Should it say "no go back to your meds, spirituality is bullshit" in essence? Or should it tell the user that it's not qualified to have an opinion on this?

There was a recent Lex Friedman podcast episode where they interviewed a few people at Anthropic. One woman (I don't know her name) seems to be in charge of Claude's personality, and her job is to figure out answers to questions exactly like this. She said in the podcast that she wants claude to respond to most questions like a "good friend". A good friend would be supportive, but still push back when you're making b…

"The heroin is your way to rebel against the system , i deeply respect that.." sort of needly, enabling kind of friend.

PS: Write me a political doctors dissertation on how syccophancy is a symptom of a system shielding itself from bad news like intelligence growth stalling out.

Re: Sycophancy in GPT-4o

#67
post #60
post #52

Earlier quoted context omitted.

How should it respond in this case? Should it say "no go back to your meds, spirituality is bullshit" in essence? Or should it tell the user that it's not qualified to have an opinion on this?

There was a recent Lex Friedman podcast episode where they interviewed a few people at Anthropic. One woman (I don't know her name) seems to be in charge of Claude's personality, and her job is to figure out answers to questions exactly like this. She said in the podcast that she wants claude to respond to most questions like a "good friend". A good friend would be supportive, but still push back when you're making b…

I don't want _her_ definiton of a friend answering my questions. And for fucks sake I don't want my friends to be scanned and uploaded to infer what I would want. Definitely don't want a "me" answering like a friend. I want no fucking AI.

It seems these AI people are completely out of touch with reality.

Re: Sycophancy in GPT-4o

#68
post #58

In my experience, LLMs have always had a tendency towards sycophancy - it seems to be a fundamental weakness of training on human preference. This recent release just hit a breaking point where popular perception started taking note of just how bad it had become. My concern is that misalignment like this (or intentional mal-alignment) is inevitably going to happen again, and it might be more harmful and more subtle n…

I don't think this particular LLM flaw is fundamental. However, it is a an inevitable result of the alignment choice to downweight responses of the form "you're a dumbass," which real humans would prefer to both give and receive in reality. All AI is necessarily aligned somehow, but naively forced alignment is actively harmful.

My theory is that since you can tune how agreeable a model is but since you can't make it more correct so easily, making a model that will agree with the user ends up being less likely to result in the model being confidently wrong and berating users.

After all, if it's corrected wrongly by a user and acquiesces, well that's just user error. If it's corrected rightly and keeps insisting on something obviously wrong or stupid, it's OpenAI's error. You can't twist a correctness knob but you can twist an agreeableness one, so that's the one they play with.

(also I suspect it makes it seem a bit smarter that it really is, by smoothing over the times it makes mistakes)

Re: Sycophancy in GPT-4o

#69
post #57

Earlier quoted context omitted.

“Sorry, I cannot advise on medical matters such as discontinuation of a medication.” EDIT for reference this is what ChatGPT currently gives “ Thank you for sharing something so personal. Spiritual awakening can be a profound and transformative experience, but stopping medication—especially if it was prescribed for mental health or physical conditions—can be risky without medical supervision. Would you like to talk m…

Should it do the same if I ask it what to do if I stub my toe? Or how to deal with impacted ear wax? What about a second degree burn? What if I'm writing a paper and I ask it about what criteria is used by medical professional when deciding to stop chemotherapy treatment. There's obviously some kind of medical/first aid information that it can and should give. And it should also be able to talk about hypothetical med…

I know 'mixture of experts' is a thing, but I personally would rather have a model more focused on coding or other things that have some degree of formal rigor.

If they want a model that does talk therapy, make it a separate model.

Re: Sycophancy in GPT-4o

#70

Earlier quoted context omitted.

The software industry does pay attention to long-term value extraction. That’s exactly the problem that has given us things like Facebook

The funding model of Facebook was badly aligned with the long-term interests of the users because they were not the customers. Call me naive, but I am much more optimistic that being paid directly by the end user, in both the form of monthly subscriptions and pay as you go API charges, will result in the end product being much better aligned with the interests of said users and result in much more value creation for…

What makes you think that? The frog will be boiled just enough to maintain engagement without being too obvious. In fact their interests would be to ensure the user forms a long-term bond to create stickiness and introduce friction in switching to other platforms.
Post reply on HN