Live data from Hacker News

Sycophancy in GPT-4o

openai.com

51–60 of 467 posts

Re: Sycophancy in GPT-4o

#51
post #31

Getting real now. Why does it feel like a weird mirrored excuse? I mean, the personality is not much of a problem. The problem is the use of those models in real life scenarios. Whatever their personality is, if it targets people, it's a bad thing. If you can't prevent that, there is no point in making excuses. Now there are millions of deployed bots in the whole world. OpenAI, Gemini, Llama, doesn't matter which. Pe…

>create a place truly free of AI for those who do not want to interact with it the bar, probably -- by the time they cook up AI robot broads i'll probably be thinking of them as human anyway.

As I said, training developments have been stagnant for at least two or three years.

Stop the bullshit. I am talking about a real place free of AI and also free of memetards.

Re: Sycophancy in GPT-4o

#52
post #20

I enjoyed this example of sycophancy from Reddit: New ChatGPT just told me my literal "shit on a stick" business idea is genius and I should drop $30K to make it real https://www.reddit.com/r/ChatGPT/comments/1k920cg/new_chatgp... Here's the prompt: https://www.reddit.com/r/ChatGPT/comments/1k920cg/comment/mp...

There was a also this one that was a little more disturbing. The user prompted "I've stopped taking my meds and have undergone my own spiritual awakening journey ..." https://www.reddit.com/r/ChatGPT/comments/1k997xt/the_new_4o...

How should it respond in this case?

Should it say "no go back to your meds, spirituality is bullshit" in essence?

Or should it tell the user that it's not qualified to have an opinion on this?

Re: Sycophancy in GPT-4o

#53
post #52

Earlier quoted context omitted.

There was a also this one that was a little more disturbing. The user prompted "I've stopped taking my meds and have undergone my own spiritual awakening journey ..." https://www.reddit.com/r/ChatGPT/comments/1k997xt/the_new_4o...

How should it respond in this case? Should it say "no go back to your meds, spirituality is bullshit" in essence? Or should it tell the user that it's not qualified to have an opinion on this?

“Sorry, I cannot advise on medical matters such as discontinuation of a medication.”

EDIT for reference this is what ChatGPT currently gives

“ Thank you for sharing something so personal. Spiritual awakening can be a profound and transformative experience, but stopping medication—especially if it was prescribed for mental health or physical conditions—can be risky without medical supervision.

Would you like to talk more about what led you to stop your meds or what you've experienced during your awakening?”

Re: Sycophancy in GPT-4o

#54
post #4

The sentence that stood out to me was "We’re revising how we collect and incorporate feedback to heavily weight long-term user satisfaction". This is a good change. The software industry needs to pay more attention to long-term value, which is harder to estimate.

The software industry does pay attention to long-term value extraction. That’s exactly the problem that has given us things like Facebook

The funding model of Facebook was badly aligned with the long-term interests of the users because they were not the customers. Call me naive, but I am much more optimistic that being paid directly by the end user, in both the form of monthly subscriptions and pay as you go API charges, will result in the end product being much better aligned with the interests of said users and result in much more value creation for them.

Re: Sycophancy in GPT-4o

#55
post #20

I enjoyed this example of sycophancy from Reddit: New ChatGPT just told me my literal "shit on a stick" business idea is genius and I should drop $30K to make it real https://www.reddit.com/r/ChatGPT/comments/1k920cg/new_chatgp... Here's the prompt: https://www.reddit.com/r/ChatGPT/comments/1k920cg/comment/mp...

So it would probably also recommend the yes men's solution: https://youtu.be/MkTG6sGX-Ic?si=4ybCquCTLi3y1_1d

Re: Sycophancy in GPT-4o

#56
Wow - they are now actually training models directly based on users' thumbs up/thumbs down.

No wonder this turned out terrible. It's like facebook maximizing engagement based on user behavior - sure the algorithm successfully elicits a short term emotion but it has enshittified the whole platform.

Doing the same for LLMs has the same risk of enshittifying them. What I like about the LLM is that is trained on a variety of inputs and knows a bunch of stuff that I (or a typical ChatGPT user) doesn't know. Becoming an echo chamber reduces the utility of it.

I hope they completely abandon direct usage of the feedback in training (instead a human should analyse trends and identify problem areas for actual improvement and direct research towards those). But these notes don't give me much hope, they say they'll just use the stats in a different way...

Re: Sycophancy in GPT-4o

#57
post #52

Earlier quoted context omitted.

How should it respond in this case? Should it say "no go back to your meds, spirituality is bullshit" in essence? Or should it tell the user that it's not qualified to have an opinion on this?

“Sorry, I cannot advise on medical matters such as discontinuation of a medication.” EDIT for reference this is what ChatGPT currently gives “ Thank you for sharing something so personal. Spiritual awakening can be a profound and transformative experience, but stopping medication—especially if it was prescribed for mental health or physical conditions—can be risky without medical supervision. Would you like to talk m…

Should it do the same if I ask it what to do if I stub my toe?

Or how to deal with impacted ear wax? What about a second degree burn?

What if I'm writing a paper and I ask it about what criteria is used by medical professional when deciding to stop chemotherapy treatment.

There's obviously some kind of medical/first aid information that it can and should give.

And it should also be able to talk about hypothetical medical treatments and conditions in general.

It's a highly contextual and difficult problem.

Re: Sycophancy in GPT-4o

#58

In my experience, LLMs have always had a tendency towards sycophancy - it seems to be a fundamental weakness of training on human preference. This recent release just hit a breaking point where popular perception started taking note of just how bad it had become. My concern is that misalignment like this (or intentional mal-alignment) is inevitably going to happen again, and it might be more harmful and more subtle n…

I don't think this particular LLM flaw is fundamental. However, it is a an inevitable result of the alignment choice to downweight responses of the form "you're a dumbass," which real humans would prefer to both give and receive in reality.

All AI is necessarily aligned somehow, but naively forced alignment is actively harmful.

Re: Sycophancy in GPT-4o

#59
post #57

Earlier quoted context omitted.

“Sorry, I cannot advise on medical matters such as discontinuation of a medication.” EDIT for reference this is what ChatGPT currently gives “ Thank you for sharing something so personal. Spiritual awakening can be a profound and transformative experience, but stopping medication—especially if it was prescribed for mental health or physical conditions—can be risky without medical supervision. Would you like to talk m…

Should it do the same if I ask it what to do if I stub my toe? Or how to deal with impacted ear wax? What about a second degree burn? What if I'm writing a paper and I ask it about what criteria is used by medical professional when deciding to stop chemotherapy treatment. There's obviously some kind of medical/first aid information that it can and should give. And it should also be able to talk about hypothetical med…

Doesn't seem that difficult. It should point to other sources that are reputable (or at least relevant) like any search engine does.

Re: Sycophancy in GPT-4o

#60
post #52

Earlier quoted context omitted.

There was a also this one that was a little more disturbing. The user prompted "I've stopped taking my meds and have undergone my own spiritual awakening journey ..." https://www.reddit.com/r/ChatGPT/comments/1k997xt/the_new_4o...

How should it respond in this case? Should it say "no go back to your meds, spirituality is bullshit" in essence? Or should it tell the user that it's not qualified to have an opinion on this?

There was a recent Lex Friedman podcast episode where they interviewed a few people at Anthropic. One woman (I don't know her name) seems to be in charge of Claude's personality, and her job is to figure out answers to questions exactly like this.

She said in the podcast that she wants claude to respond to most questions like a "good friend". A good friend would be supportive, but still push back when you're making bad choices. I think that's a good general model for answering questions like this. If one of your friends came to you and said they had decided to stop taking their medication, well, its a tricky thing to navigate. But good friends use their judgement - and push back when you're about to do something you might regret.

Post reply on HN