Live data from Hacker News

Sycophancy in GPT-4o

openai.com

221–230 of 467 posts

Re: Sycophancy in GPT-4o

#222

Tragically, ChatGPT might be the only "one" who sycophants the user. From students to workforce, who is getting compliments and encouragement that they are doing well. In a not so far future dystopia, we might have kids who remember that the only kind and encourage soul in their childhood was something without a soul.

Fantastic insight, thanks!

Re: Sycophancy in GPT-4o

#223
post #176

As an engineer, I need AIs to tell me when something is wrong or outright stupid. I'm not seeking validation, I want solutions that work. 4o was unusable because of this, very glad to see OpenAI walk back on it and recognise their mistake. Hopefully they learned from this and won't repeat the same errors, especially considering the devastating effects of unleashing THE yes-man on people who do not have the mental cap…

It's a recipe for disaster.

Frankly, I think it's genuinely dangerous.

Re: Sycophancy in GPT-4o

#224

Earlier quoted context omitted.

Only AI enthusiasts know about Grok, and only some dedicated subset of fans are advocating for it. Meanwhile even my 97 year old grandfather heard about ChatGPT.

This. Only on HN does ChatGPT somehow fear losing customers to Grok. Until Grok works out how to market to my mother, or at least make my mother aware that it exists, taking ChatGPT customers ain't happening.

They are cargoculting. Almost literally. It's MO for Musk companies.

They might call it open discussion and startup style rapid iteration approach, but they aren't getting it. Their interpretation of it is just collective hallucination under assumption that adults come to change diapers.

Re: Sycophancy in GPT-4o

#225

Earlier quoted context omitted.

For us habitual users of em-dashes, it is saddening to have to think twice about using them lest someone think we are using an LLM to write…

Most keyboards don't have an em-dash key, so what do you expect?

I also use em-dash regularly. In Microsoft Outlook and Microsoft Word, when you type double dash, then space, it will be converted to an em-dash. This is how most normies type an em-dash.

Re: Sycophancy in GPT-4o

#226

Earlier quoted context omitted.

One of the biggest tells.

For us habitual users of em-dashes, it is saddening to have to think twice about using them lest someone think we are using an LLM to write…

I use the en-dash (Alt+0150) instead of the em.

The en-dash and the em-dash are interchangeable in Finnish. The shorter form has more "inoffensive" look-and-feel and maybe that's why it's used more often here.

Now that I think of it, I don't seem to remember the alt code of the em-dash...

Re: Sycophancy in GPT-4o

#227

Earlier quoted context omitted.

I don't think they were imitating grok, they were aiming to improve retention but it backfired and ended up being too on-the-nose (if they had a choice they wouldn't wanted it to be this obvious). Grok has it's own "default voice" which I sort of dislike, it tries too hard to seem "hip" for lack of a better word.

> it tries too hard to seem "hip" for lack of a better word. Reminds me of someone.

However, I hope it gives better advice than the someone you're thinking of. But Grok's training data is probably more balanced than that used by you-know-who (which seems to be "all of rightwing X")...

Re: Sycophancy in GPT-4o

#228

Earlier quoted context omitted.

Only AI enthusiasts know about Grok, and only some dedicated subset of fans are advocating for it. Meanwhile even my 97 year old grandfather heard about ChatGPT.

First mover advantage. This won't change. Same as Xerox vs photocopy. I use Grok myself but talk about ChatGPT is my blog articles when I write something related to LLM.

That's... not really an advertisement for your blog, is it?

Re: Sycophancy in GPT-4o

#229

It's worth noting that one of the fixes OpenAI employed to get ChatGPT to stop being sycophantic is to simply to edit the system prompt to include the phrase "avoid ungrounded or sycophantic flattery": https://simonwillison.net/2025/Apr/29/chatgpt-sycophancy-pro... I personally never use the ChatGPT webapp or any other chatbot webapps — instead using the APIs directly — because being able to control the system prompt…

You can bypass the system prompt by using the API? I thought part of the "safety" of LLMs was implemented with the system prompt. Does that mean it's easier to get unsafe answers by using the API instead of the GUI?

Re: Sycophancy in GPT-4o

#230
post #200

In my experience, LLMs have always had a tendency towards sycophancy - it seems to be a fundamental weakness of training on human preference. This recent release just hit a breaking point where popular perception started taking note of just how bad it had become. My concern is that misalignment like this (or intentional mal-alignment) is inevitably going to happen again, and it might be more harmful and more subtle n…

> In my experience, LLMs have always had a tendency towards sycophancy The very early ones (maybe GPT 3.0?) sure didn't. You'd show them they were wrong, and they'd say something that implied that OK maybe you were right, but they weren't so sure; or that their original mistake was your fault somehow.

Were those trained using RLHF? IIRC the earliest models were just using SFT for instruction following.

Like the GP said, I think this is fundamentally a problem of training on human preference feedback. You end up with a model that produces things that cater to human preferences, which (necessarily?) includes the degenerate case of sycophancy.

Post reply on HN