Live data from Hacker News

Expanding on what we missed with sycophancy

openai.com

21–30 of 297 posts

Re: Expanding on what we missed with sycophancy

#21
post #3

OpenAI mentions the new memory features as a partial cause. My theory as a imperative/functional programmer is that those features added global state to prompts that didn't have it before leading to unpredictability and instabilty. Prompts went from stateless to stateful. As GPT 4o put it: 1. State introduces non-determinism across sessions 2. Memory + sycophancy is a feedback loop 3. Memory acts as a shadow prompt m…

It is. If you start a fresh chat, turn on advanced voice, and just make any random sound like snapping your fingers it will just randomly pick up as if you’re continuing some other chat with no context (on the user side). I honestly really dislike that it considers all my previous interactions because I typically used new chats as a way to get it out of context ruts.

I don't like the change either. At the least it should be an option you can configure. But, can you use a "temporary" chat to ignore your other chats as a workaround?

Re: Expanding on what we missed with sycophancy

#22

My layman’s view is that this issue was primarily due to the fact that 4o is no longer their flagship model. Similar to the Ford Mustang, much of the performance efforts are on the higher trims, while the base trims just get larger and louder engines, because that’s what users want. With presumably everyone at OpenAI primarily using the newest models (o3), the updates to the base user model have been further automate…

[deleted]

Re: Expanding on what we missed with sycophancy

#23

I found the recent sycophancy a bit annoying when trying to diagnose and solve coding problems. First it would waste time praising your intelligence for asking the question before getting to the answer. But more annoyingly if I asked "I am encountering X issue, could Y be the cause" or "could Y be a solution", the response would nearly always be "yes, exactly, it's Y" even when it wasn't the case. I guess part of the…

For many people ChatGPT is already the smartest relationship they have in their lives, not sure how long we have until it’s the most fulfilling. On the upside it is plausible that ChatGPT can get to a state where it can act as a good therapist and help helpless who otherwise would not get help.

I am more regularly finding myself in discussions where the other person believes they’re right because they have ChatGPT in their corner.

I think most smart people overestimate the intelligence of others for a variety of reasons so they overestimate what it would take for a LLM to beat the output of an average person.

Re: Expanding on what we missed with sycophancy

#24

Earlier quoted context omitted.

It is. If you start a fresh chat, turn on advanced voice, and just make any random sound like snapping your fingers it will just randomly pick up as if you’re continuing some other chat with no context (on the user side). I honestly really dislike that it considers all my previous interactions because I typically used new chats as a way to get it out of context ruts.

I don't like the change either. At the least it should be an option you can configure. But, can you use a "temporary" chat to ignore your other chats as a workaround?

[deleted]

Re: Expanding on what we missed with sycophancy

#25

Earlier quoted context omitted.

It is. If you start a fresh chat, turn on advanced voice, and just make any random sound like snapping your fingers it will just randomly pick up as if you’re continuing some other chat with no context (on the user side). I honestly really dislike that it considers all my previous interactions because I typically used new chats as a way to get it out of context ruts.

I don't like the change either. At the least it should be an option you can configure. But, can you use a "temporary" chat to ignore your other chats as a workaround?

I had a discussion with GPT 4o about the memory system. I'd don't know if any of this is made up but it's a start for further research

- Memory in settings is configurable. It is visible and can be edited.

- Memory from global chat history is not configurable. Think of it as a system cache.

- Both memory systems can be turned off

- Chats in Projects do not use the global chat history. They are isolated.

- Chats in Projects do use settings memory but that can be turned off.

Re: Expanding on what we missed with sycophancy

#26

My layman’s view is that this issue was primarily due to the fact that 4o is no longer their flagship model. Similar to the Ford Mustang, much of the performance efforts are on the higher trims, while the base trims just get larger and louder engines, because that’s what users want. With presumably everyone at OpenAI primarily using the newest models (o3), the updates to the base user model have been further automate…

Anecdotally, there was also a strong correlation between high-sycophancy and high-quality that cooked up recently. I was voting for equations/tables rather than overwrought blocks of descriptive text, which I am pretty comfortable defending as an orthogonal concern, but the "sycophancy gene" always landed on the same side as the equations/tables for whatever reason. I'm pretty sure this isn't an intrinsic connection…

[deleted]

Re: Expanding on what we missed with sycophancy

#27

My layman’s view is that this issue was primarily due to the fact that 4o is no longer their flagship model. Similar to the Ford Mustang, much of the performance efforts are on the higher trims, while the base trims just get larger and louder engines, because that’s what users want. With presumably everyone at OpenAI primarily using the newest models (o3), the updates to the base user model have been further automate…

What's a "trim" in this context?

Re: Expanding on what we missed with sycophancy

#28
post #27

My layman’s view is that this issue was primarily due to the fact that 4o is no longer their flagship model. Similar to the Ford Mustang, much of the performance efforts are on the higher trims, while the base trims just get larger and louder engines, because that’s what users want. With presumably everyone at OpenAI primarily using the newest models (o3), the updates to the base user model have been further automate…

What's a "trim" in this context?

https://en.wikipedia.org/wiki/Automotive_trim_level

Re: Expanding on what we missed with sycophancy

#30
post #13

Earlier quoted context omitted.

Are you sure that was real? I thought it was an made up example of the problems with the update

It didn't matter to me if it was real, because I believe that there are edge cases where it could happen and that warrented a shutdown and pullback. The sychophant will be back because they accidentally stumbled upon an engagement manager's dream machine.

It kind of does matter if it's real, because in my experience this is something OpenAI has thought about a lot, and added significant protections to address exactly this class of issue.

Throwing out strawman hypotheticals is just going to confuse the public debate over what protections need to be prioritized.

Post reply on HN