Live data from Hacker News

Expanding on what we missed with sycophancy

openai.com

41–50 of 297 posts

Re: Expanding on what we missed with sycophancy

#41

I find it disappointing that openai doesn't really mention anything here along the lines of having an accurate model of reality. That's really what the problem is with sycophancy, it encourages people to detach themselves from what reality is. Like, it seems like they are saying their "vibe check" didn't check vibes enough.

This is such an interesting question though! It seems to bring to the fore a lot of deeper, philosophical things like if there even IS such a thing as objective reality or objective context within which the AI should be operating. From training data, there might be some generalizations that are carried across all contexts, but that starts to not be applicable when person A with a college degree says they want to start business x versus person B without said degree who also wants to start business x, how does the model properly reconcile the context of the general advise and each asker’s unique circumstances? Does it ask an infinite list of probing questions before answering? It gets into much the same problems as issues of advise among people.

Plus, things get even harder when it comes to even less quantifiable contexts like mental health and relationships.

In all, I am not saying there isnt some approximated and usable “objective” reality, just that it starts to break down when it gets to the individual and that is where openai is failing by over-emphasizing reflective behavior in the absence if actual data about the user.

Re: Expanding on what we missed with sycophancy

#43
post #42

That doesn't make any sense to me. Seems like you're trying to blame one LLM revision for something that went wrong. It oozes a smell of unaccountability. Thus, unaligned. From tech to public relations.

Except that's literally how LLMs work. Small changes to the prompt or training can greatly affect its output.

Re: Expanding on what we missed with sycophancy

#45
post #42

That doesn't make any sense to me. Seems like you're trying to blame one LLM revision for something that went wrong. It oozes a smell of unaccountability. Thus, unaligned. From tech to public relations.

Except that's literally how LLMs work. Small changes to the prompt or training can greatly affect its output.

[flagged]

Re: Expanding on what we missed with sycophancy

#46
post #42

That doesn't make any sense to me. Seems like you're trying to blame one LLM revision for something that went wrong. It oozes a smell of unaccountability. Thus, unaligned. From tech to public relations.

I can totally believe that they deployed it because internal metrics looked good.

Re: Expanding on what we missed with sycophancy

#48
post #45

Earlier quoted context omitted.

Except that's literally how LLMs work. Small changes to the prompt or training can greatly affect its output.

[flagged]

I legitimately don't understand what your point is. Are you saying that they actually did this intentionally?

Re: Expanding on what we missed with sycophancy

#49

My most cynical take is that this is OpenAI's Conway's Law problem, and it reflects the structure and sycophancy of the organization broadly all the way up to sama. That company has seen a lot of talent attrition over the last year—the type of talent that would have pushed back against outcomes like this. I think we'll continue to see this kind of thing play out for a while. Oh GPT, you're just like your father!

You may be thinking of Conway's "how committees invent" paper.

Re: Expanding on what we missed with sycophancy

#50
This is not truly solvable. There is an extremely strong outer loop of optimization operating here: we want it.

We will use models that make us feel good over models that don't make us feel good.

This one was a little too ham-fisted (at least, for the sensibilities of people in our media bubble; though I suspect there is also an enormous mass of people for whom it was not), so they turned it down a bit. Later iterations will be subtler, and better at picking up the exact level and type of sycophancy that makes whoever it's talking to unsuspiciously feel good (feel right, feel smart, feel understood, etc).

It'll eventually disappear, to you, as it's dialed in, to you.

This may be the medium-term fate of both LLMs and humans, only resolved when the humans wither away.

Post reply on HN