OpenAI mentions the new memory features as a partial cause. My theory as a imperative/functional programmer is that those features added global state to prompts that didn't have it before leading to unpredictability and instabilty. Prompts went from stateless to stateful. As GPT 4o put it: 1. State introduces non-determinism across sessions 2. Memory + sycophancy is a feedback loop 3. Memory acts as a shadow prompt m…
It is. If you start a fresh chat, turn on advanced voice, and just make any random sound like snapping your fingers it will just randomly pick up as if you’re continuing some other chat with no context (on the user side). I honestly really dislike that it considers all my previous interactions because I typically used new chats as a way to get it out of context ruts.
Expanding on what we missed with sycophancy
21–30 of 297 posts
Re: Expanding on what we missed with sycophancy
#22My layman’s view is that this issue was primarily due to the fact that 4o is no longer their flagship model. Similar to the Ford Mustang, much of the performance efforts are on the higher trims, while the base trims just get larger and louder engines, because that’s what users want. With presumably everyone at OpenAI primarily using the newest models (o3), the updates to the base user model have been further automate…
Re: Expanding on what we missed with sycophancy
#23I found the recent sycophancy a bit annoying when trying to diagnose and solve coding problems. First it would waste time praising your intelligence for asking the question before getting to the answer. But more annoyingly if I asked "I am encountering X issue, could Y be the cause" or "could Y be a solution", the response would nearly always be "yes, exactly, it's Y" even when it wasn't the case. I guess part of the…
I am more regularly finding myself in discussions where the other person believes they’re right because they have ChatGPT in their corner.
I think most smart people overestimate the intelligence of others for a variety of reasons so they overestimate what it would take for a LLM to beat the output of an average person.
Re: Expanding on what we missed with sycophancy
#24Earlier quoted context omitted.
It is. If you start a fresh chat, turn on advanced voice, and just make any random sound like snapping your fingers it will just randomly pick up as if you’re continuing some other chat with no context (on the user side). I honestly really dislike that it considers all my previous interactions because I typically used new chats as a way to get it out of context ruts.
I don't like the change either. At the least it should be an option you can configure. But, can you use a "temporary" chat to ignore your other chats as a workaround?
Re: Expanding on what we missed with sycophancy
#25Earlier quoted context omitted.
It is. If you start a fresh chat, turn on advanced voice, and just make any random sound like snapping your fingers it will just randomly pick up as if you’re continuing some other chat with no context (on the user side). I honestly really dislike that it considers all my previous interactions because I typically used new chats as a way to get it out of context ruts.
I don't like the change either. At the least it should be an option you can configure. But, can you use a "temporary" chat to ignore your other chats as a workaround?
- Memory in settings is configurable. It is visible and can be edited.
- Memory from global chat history is not configurable. Think of it as a system cache.
- Both memory systems can be turned off
- Chats in Projects do not use the global chat history. They are isolated.
- Chats in Projects do use settings memory but that can be turned off.
Re: Expanding on what we missed with sycophancy
#26My layman’s view is that this issue was primarily due to the fact that 4o is no longer their flagship model. Similar to the Ford Mustang, much of the performance efforts are on the higher trims, while the base trims just get larger and louder engines, because that’s what users want. With presumably everyone at OpenAI primarily using the newest models (o3), the updates to the base user model have been further automate…
Anecdotally, there was also a strong correlation between high-sycophancy and high-quality that cooked up recently. I was voting for equations/tables rather than overwrought blocks of descriptive text, which I am pretty comfortable defending as an orthogonal concern, but the "sycophancy gene" always landed on the same side as the equations/tables for whatever reason. I'm pretty sure this isn't an intrinsic connection…
Re: Expanding on what we missed with sycophancy
#27My layman’s view is that this issue was primarily due to the fact that 4o is no longer their flagship model. Similar to the Ford Mustang, much of the performance efforts are on the higher trims, while the base trims just get larger and louder engines, because that’s what users want. With presumably everyone at OpenAI primarily using the newest models (o3), the updates to the base user model have been further automate…
Re: Expanding on what we missed with sycophancy
#28My layman’s view is that this issue was primarily due to the fact that 4o is no longer their flagship model. Similar to the Ford Mustang, much of the performance efforts are on the higher trims, while the base trims just get larger and louder engines, because that’s what users want. With presumably everyone at OpenAI primarily using the newest models (o3), the updates to the base user model have been further automate…
What's a "trim" in this context?
Re: Expanding on what we missed with sycophancy
#29Re: Expanding on what we missed with sycophancy
#30Earlier quoted context omitted.
Are you sure that was real? I thought it was an made up example of the problems with the update
It didn't matter to me if it was real, because I believe that there are edge cases where it could happen and that warrented a shutdown and pullback. The sychophant will be back because they accidentally stumbled upon an engagement manager's dream machine.
Throwing out strawman hypotheticals is just going to confuse the public debate over what protections need to be prioritized.