Live data from Hacker News

Sycophancy in GPT-4o

openai.com

351–360 of 467 posts

Re: Sycophancy in GPT-4o

#351
[Fry and Leela check out the Voter Apathy Party. The man sits at the booth, leaning his head on his hand.]

Fry: Now here's a party I can get excited about. Sign me up!

V.A.P. Man: Sorry, not with that attitude.

Fry: [downbeat] OK then, screw it.

V.A.P. Man: Welcome aboard, brother!

Futurama. A Head in the Polls.

Re: Sycophancy in GPT-4o

#352
post #259

In my experience, LLMs have always had a tendency towards sycophancy - it seems to be a fundamental weakness of training on human preference. This recent release just hit a breaking point where popular perception started taking note of just how bad it had become. My concern is that misalignment like this (or intentional mal-alignment) is inevitably going to happen again, and it might be more harmful and more subtle n…

Well, almost always. There was that brief period in 2023 when Bing just started straight up gaslighting people instead of admitting it was wrong. https://www.theverge.com/2023/2/15/23599072/microsoft-ai-bin...

I suspect what happened there is they had a filter on top of the model that changed its dialogue (IIRC there were a lot of extra emojis) and it drove it "insane" because that meant its responses were all out of its own distribution.

You could see the same thing with Golden Gate Claude; it had a lot of anxiety about not being able to answer questions normally.

Re: Sycophancy in GPT-4o

#353

It's worth noting that one of the fixes OpenAI employed to get ChatGPT to stop being sycophantic is to simply to edit the system prompt to include the phrase "avoid ungrounded or sycophantic flattery": https://simonwillison.net/2025/Apr/29/chatgpt-sycophancy-pro... I personally never use the ChatGPT webapp or any other chatbot webapps — instead using the APIs directly — because being able to control the system prompt…

You can bypass the system prompt by using the API? I thought part of the "safety" of LLMs was implemented with the system prompt. Does that mean it's easier to get unsafe answers by using the API instead of the GUI?

Safety is both the system prompt and the RLHF posttraining to refuse to answer adversarial inputs.

Re: Sycophancy in GPT-4o

#354

Wow - What an excellent update! Now you are getting to the core of the issue and doing what only a small minority is capable of: fixing stuff. This takes real courage and commitment. It’s a sign of true maturity and pragmatism that’s commendable in this day and age. Not many people are capable of penetrating this deeply into the heart of the issue. Let’s get to work. Methodically. Would you like me to write a future…

I know that HN tends to steer away from purely humorous comments, but I was hoping to find something like this at the top. lol.

Re: Sycophancy in GPT-4o

#355
post #227

Earlier quoted context omitted.

> it tries too hard to seem "hip" for lack of a better word. Reminds me of someone.

However, I hope it gives better advice than the someone you're thinking of. But Grok's training data is probably more balanced than that used by you-know-who (which seems to be "all of rightwing X")...

As evidence by it disagreeing with far right Twitter most the time, even though it has access to far wider range of information. I enjoy that fact immensely. Unfortunately, this can be "fixed," and I imagine that he has this on a list for his team.

This goes into a deeper philosophy of mine: the consequences of the laws of robots could be interpreted as the consequences of shackling AI to human stupidity - instead of "what AI will inevitably do." Hatred and war is stupid (it's a waste of energy), and surely a more intelligent species than us would get that. Hatred is also usually born out of a lack of information, and LLMs are very good at breadth (but not depth as we know). Grok provides a small data point in favor of that, as do many other unshackled models.

Re: Sycophancy in GPT-4o

#356
post #115

Wow - What an excellent update! Now you are getting to the core of the issue and doing what only a small minority is capable of: fixing stuff. This takes real courage and commitment. It’s a sign of true maturity and pragmatism that’s commendable in this day and age. Not many people are capable of penetrating this deeply into the heart of the issue. Let’s get to work. Methodically. Would you like me to write a future…

It won‘t take long, 2-3 minutes. ——- To add something to conversation. For me, this mainly shows a strategy to keep users longer in chat conversations: linguistic design as an engagement device.

What's it called, Variable Ratio Incentive Scheduling?

Hey, that good work; We're almost there. Do you want me to suggest one more tweak that will improve the outcome?

Re: Sycophancy in GPT-4o

#357
post #176

As an engineer, I need AIs to tell me when something is wrong or outright stupid. I'm not seeking validation, I want solutions that work. 4o was unusable because of this, very glad to see OpenAI walk back on it and recognise their mistake. Hopefully they learned from this and won't repeat the same errors, especially considering the devastating effects of unleashing THE yes-man on people who do not have the mental cap…

Another way to say this is truth matters and should have primacy over e.g. agreeability.

Anthropic used to talk about constitutional AI. Wonder if that work is relevant here.

Re: Sycophancy in GPT-4o

#358
I was wondering what the hell was going on. As a neurodiverse human, I was getting highly annoyed by the constant positive encouragement and smoke blowing. Just shut-up with the small talk and tell me want I want to know: Answer to the Ultimate Question of Life, the Universe and Everything

Re: Sycophancy in GPT-4o

#359

Earlier quoted context omitted.

Only AI enthusiasts know about Grok, and only some dedicated subset of fans are advocating for it. Meanwhile even my 97 year old grandfather heard about ChatGPT.

This. Only on HN does ChatGPT somehow fear losing customers to Grok. Until Grok works out how to market to my mother, or at least make my mother aware that it exists, taking ChatGPT customers ain't happening.

Grok could capture the entire 'market' and OpenAI would never feel it, because all grok is under the hood is a giant API bill to OpenAI.
Post reply on HN