Live data from Hacker News

Sycophancy in GPT-4o

openai.com

81–90 of 467 posts

Re: Sycophancy in GPT-4o

#81
There has been this weird trend going around to use ChatGPT to "red team" or "find critical life flaws" or "understand what is holding me back" going around - I've read a few of them and on one hand I really like it encouraging people to "be their best them", on the other... king of spain is just genuinely out of reach of some.

Re: Sycophancy in GPT-4o

#82
post #74
post #24

Earlier quoted context omitted.

I also started by using APIs directly, but I've found that Google's AI Studio offers a good mix of the chatbot webapps and system prompt tweakability.

I find it maddening that AI Studio doesn't have a way to save the system prompt as a default.

On the top right click the save icon

Re: Sycophancy in GPT-4o

#84
What should be the solution here? There's a thing that, despite how much it may mimic humans, isn't human, and doesn't operate on the same axes. The current AI neither is nor isn't [any particular personality trait]. We're applying human moral and value judgments to something that doesn't, can't, hold any morals or values.

There's an argument to be made for, don't use the thing for which it wasn't intended. There's another argument to be made for, the creators of the thing should be held to some baseline of harm prevention; if a thing can't be done safely, then it shouldn't be done at all.

Re: Sycophancy in GPT-4o

#85

At the bottom of the page is a "Ask GPT ..." field which I thought allows users to ask questions about the page, but it just opens up ChatGPT. Missed opportunity.

no, its sensible because you need auth wall for that or it will be abused to bits

Re: Sycophancy in GPT-4o

#87

Earlier quoted context omitted.

I'm not sure how this problem can be solved. How do you test a system with emergent properties of this degree that whose behavior is dependent on existing memory of customer chats in production?

Using prompts know to be problematic? Some sort of... Voight-Kampff test for LLMs?

I doubt it's that simple. What about memories running in prod? What about explicit user instructions? What about subtle changes in prompts? What happens when a bad release poisons memories?

The problem space is massive and is growing rapidly, people are finding new ways to talk to LLMs all the time

Re: Sycophancy in GPT-4o

#88
On a different note, does that mean that specifying "4o" doesn't always get you the same model? If you pin a particular operation to use "4o", they could still swap the model out from under you, and maybe the divergence in behavior breaks your usage?

Re: Sycophancy in GPT-4o

#89

On a different note, does that mean that specifying "4o" doesn't always get you the same model? If you pin a particular operation to use "4o", they could still swap the model out from under you, and maybe the divergence in behavior breaks your usage?

If you look in the API there are several flavors of 4o that behave fairly differently.

Re: Sycophancy in GPT-4o

#90
They are talking about how their thumbs up / thumbs down signal were applied incorrectly, because they dont represent what they thought they measure.

If only there was a way to gather feedback in a more verbose way, where user can specify what he liked and didnt about the answer, and extract that sentiment at scale...

Post reply on HN