Live data from Hacker News

Sycophancy in GPT-4o

openai.com

181–190 of 467 posts

Re: Sycophancy in GPT-4o

#181
I like they learned these adjustments didn't 'work'. My concern is what if OpenAI is to do subtle A/B testing based on previous interactions and optimize interactions based on users personality/mood? Maybe not telling you 'shit on a stick' is awesome idea, but being able to steer you towards a conclusion sort of like [1].

[1] https://www.newscientist.com/article/2478336-reddit-users-we...

Re: Sycophancy in GPT-4o

#182
post #20

I enjoyed this example of sycophancy from Reddit: New ChatGPT just told me my literal "shit on a stick" business idea is genius and I should drop $30K to make it real https://www.reddit.com/r/ChatGPT/comments/1k920cg/new_chatgp... Here's the prompt: https://www.reddit.com/r/ChatGPT/comments/1k920cg/comment/mp...

There was a also this one that was a little more disturbing. The user prompted "I've stopped taking my meds and have undergone my own spiritual awakening journey ..." https://www.reddit.com/r/ChatGPT/comments/1k997xt/the_new_4o...

there was one on twitter where people would talk like they had Intelligence attribute set to 1 and GPT would praise them for being so smart

Re: Sycophancy in GPT-4o

#183

Earlier quoted context omitted.

I was about to roast you until I realized this had to be satire given the situation, haha. They tried to imitate grok with a cheaply made system prompt, it had an uncanny effect, likely because it was built on a shaky foundation. And now they are trying to save face before they lose customers to Grok 3.5 which is releasing in beta early next week.

Only AI enthusiasts know about Grok, and only some dedicated subset of fans are advocating for it. Meanwhile even my 97 year old grandfather heard about ChatGPT.

This.

Only on HN does ChatGPT somehow fear losing customers to Grok. Until Grok works out how to market to my mother, or at least make my mother aware that it exists, taking ChatGPT customers ain't happening.

Re: Sycophancy in GPT-4o

#185

We should be loudly demanding transparency. If you're auto-opted into the latest model revision, you don't know what you're getting day-to-day. A hammer behaves the same way every time you pick it up; why shouldn't LLMs? Because convenience. Convenience features are bad news if you need to be as a tool. Luckily you can still disable ChatGPT memory. Latent Space breaks it down well - the "tool" (Anton) vs. "magic" (Cl…

> why shouldn't LLMs

Because they're non-deterministic.

Re: Sycophancy in GPT-4o

#186
I wanted to see how far it will go. I started with asking it to simple test app. It said it is a great idea. And asked me if I want to do market analysis. I came back later and asked it to do a TAM analysis. It said $2-20B. Then it asked if it can make a one page investor pitch. I said ok, go ahead. Then it asked if I want a detailed slide deck. After making the deck it asked if I want a keynote file for the deck.

All this while I was thinking this is more dangerous than instagram. Instagram only sent me to the gym and to touristic places and made me buy some plastic. ChatGPT wants me to be a tech bro and speed track the Billion dollar net worth.

Re: Sycophancy in GPT-4o

#187
> The update we removed was overly flattering or agreeable—often described as sycophantic.

> We have rolled back last week’s GPT‑4o update in ChatGPT so people are now using an earlier version with more balanced behavior.

I thought every major LLM was extremely sycophantic. Did GPT-4o do it more than usual?

Re: Sycophancy in GPT-4o

#189
post #125

Earlier quoted context omitted.

I do think the blog post has a sycophantic vibe too. Not sure if that‘s intended.

It also has an em-dash

A remarkable insight—often associated with individuals of above-average cognitive capabilities.

While the use of the em-dash has recently been associated with AI you might offend real people using it organically—often writers and literary critics.

To conclude it’s best to be hesitant and, for now, refrain from judging prematurely.

Would you like me to elaborate on this issue or do you want to discuss some related topic?

Re: Sycophancy in GPT-4o

#190

Earlier quoted context omitted.

One of the biggest tells.

For us habitual users of em-dashes, it is saddening to have to think twice about using them lest someone think we are using an LLM to write…

Does it really matter though? I just focus on the point someone is trying to make, not on the tools they use to make it.
Post reply on HN