I found the recent sycophancy a bit annoying when trying to diagnose and solve coding problems. First it would waste time praising your intelligence for asking the question before getting to the answer. But more annoyingly if I asked "I am encountering X issue, could Y be the cause" or "could Y be a solution", the response would nearly always be "yes, exactly, it's Y" even when it wasn't the case. I guess part of the…
For many people ChatGPT is already the smartest relationship they have in their lives, not sure how long we have until it’s the most fulfilling. On the upside it is plausible that ChatGPT can get to a state where it can act as a good therapist and help helpless who otherwise would not get help. I am more regularly finding myself in discussions where the other person believes they’re right because they have ChatGPT in…
Expanding on what we missed with sycophancy
31–40 of 297 posts
Re: Expanding on what we missed with sycophancy
#32My layman’s view is that this issue was primarily due to the fact that 4o is no longer their flagship model. Similar to the Ford Mustang, much of the performance efforts are on the higher trims, while the base trims just get larger and louder engines, because that’s what users want. With presumably everyone at OpenAI primarily using the newest models (o3), the updates to the base user model have been further automate…
Re: Expanding on what we missed with sycophancy
#33Re: Expanding on what we missed with sycophancy
#34Earlier quoted context omitted.
It didn't matter to me if it was real, because I believe that there are edge cases where it could happen and that warrented a shutdown and pullback. The sychophant will be back because they accidentally stumbled upon an engagement manager's dream machine.
It kind of does matter if it's real, because in my experience this is something OpenAI has thought about a lot, and added significant protections to address exactly this class of issue. Throwing out strawman hypotheticals is just going to confuse the public debate over what protections need to be prioritized.
Seems like asserting hypothetical "significant protections to address exactly this class of issue" does the same thing though?
Re: Expanding on what we missed with sycophancy
#35I am really curious what their testing suite looks like. How do you test for sycophants?
The now rolled back model failed spectacularly on this test
Re: Expanding on what we missed with sycophancy
#36Re: Expanding on what we missed with sycophancy
#37I am really curious what their testing suite looks like. How do you test for sycophants?
Re: Expanding on what we missed with sycophancy
#38Re: Expanding on what we missed with sycophancy
#39Earlier quoted context omitted.
It is. If you start a fresh chat, turn on advanced voice, and just make any random sound like snapping your fingers it will just randomly pick up as if you’re continuing some other chat with no context (on the user side). I honestly really dislike that it considers all my previous interactions because I typically used new chats as a way to get it out of context ruts.
I don't like the change either. At the least it should be an option you can configure. But, can you use a "temporary" chat to ignore your other chats as a workaround?
Re: Expanding on what we missed with sycophancy
#40Earlier quoted context omitted.
For many people ChatGPT is already the smartest relationship they have in their lives, not sure how long we have until it’s the most fulfilling. On the upside it is plausible that ChatGPT can get to a state where it can act as a good therapist and help helpless who otherwise would not get help. I am more regularly finding myself in discussions where the other person believes they’re right because they have ChatGPT in…
I think that most smart people underestimate the complexity of fields they aren’t in. ChatGPT may be able to replace a psychology listicle, but it has no affect or ability to read, respond, and intervene or redirect like a human can.
I don’t see there being an insurmountable barrier that would prevent LLMs from doing the things you suggest it cannot. So even assuming you are correct for now I would suggest that LLMs will improve.
My estimations don’t come from my assumption that other people’s jobs are easy, they come from doing applied research in behavioral analytics on mountains of data in rather large data centers.