Live data from Hacker News

Expanding on what we missed with sycophancy

openai.com

141–150 of 297 posts

Re: Expanding on what we missed with sycophancy

#141

Earlier quoted context omitted.

I had a discussion with GPT 4o about the memory system. I'd don't know if any of this is made up but it's a start for further research - Memory in settings is configurable. It is visible and can be edited. - Memory from global chat history is not configurable. Think of it as a system cache. - Both memory systems can be turned off - Chats in Projects do not use the global chat history. They are isolated. - Chats in Pr…

I assume this is being downvoted because I said I ran it by GPT 4o. I don't know how to credit AI without giving the impression that I'm outsourcing my thinking to it

I didn't downvote but it would be because of the "I'd don't know if any of this is made up" — if you said "GPT said this, and I've verified it to be correct", that's valuable information, even it came from a language model. But otherwise (if you didn't verify), there's not much value in the post, it's basically "here is some random plausible text" and plausibly incorrect is worse than nothing.

Re: Expanding on what we missed with sycophancy

#142
The side-by-side comparisons are not a good signal because the models vary across multiple dimensions, but the user isn't given the option to indicate the dimension on which they're scoring the model.

The recent side-by-side comparisons presented a more accurate model that communicates poorly vs a less accurate model with slightly better communication.

Re: Expanding on what we missed with sycophancy

#143
post #135
post #114

Earlier quoted context omitted.

>We can argue over whether or not it's "real" empathy There's nothing to argue about, it's unambiguously not real empathy. Empathy from a human exists in a much broader context of past and future interactions. One reason human empathy is nice is because it is often followed up with actions. Friends who care about you will help you out in material ways when you need it. Even strangers will. Someone who sees a person s…

>There's nothing to argue about, it's unambiguously not real empathy I think if a person can't tell the difference between empathy from a human vs empathy from a chatbot, it's a difference without a distinction If it activates the same neural pathways, and has the same results, then I think the mind doesn't care >One reason human empathy is nice is because it is often followed up with actions. Friends who care about…

>If it activates the same neural pathways, and has the same results, then I think the mind doesn't care

Boiling it down to neural signals is a risky approach, imo. There are innumerable differences between these interactions. This isn't me saying interactions are inherently dangerous if artificial empathy is baked in, but equating them to real empathy is.

Understanding those differences is critical, especially in a world of both deliberately bad actors and those who will destroy lives in the pursuit of profit by normalizing replacements for human connections.

Re: Expanding on what we missed with sycophancy

#144
The thing that annoys me the most is when I ask it to generate some code - actually no, most often than not I don't even ask it to generate code, but ask some vaguely related programming question - to which it replies with complete listing of code (didn't ask for it, but alas).

Then I fix the code and tell it all the mistakes it has. And then it does a 180 in tone, wherein - it starts talking as if I wrote the code in the first place with - "yeah, obviously that wouldn't work, so I fixed the issues in your code" and acts like a person trying to save face and present the bugs it fixed as if the buggy code was written by me all along.

That really gets me livid. LOL

Re: Expanding on what we missed with sycophancy

#145

My layman’s view is that this issue was primarily due to the fact that 4o is no longer their flagship model. Similar to the Ford Mustang, much of the performance efforts are on the higher trims, while the base trims just get larger and louder engines, because that’s what users want. With presumably everyone at OpenAI primarily using the newest models (o3), the updates to the base user model have been further automate…

I’ve been using the 4.5 preview a lot, and it can also have a bit of a sycophantic streak, but being a larger and more intelligent model, I think it applies more nuance.

Watching this controversy, I wondered if they perhaps tried to distill 4.5’s personality into a model that is just too small to pull it off.

Re: Expanding on what we missed with sycophancy

#146

Earlier quoted context omitted.

I had a discussion with GPT 4o about the memory system. I'd don't know if any of this is made up but it's a start for further research - Memory in settings is configurable. It is visible and can be edited. - Memory from global chat history is not configurable. Think of it as a system cache. - Both memory systems can be turned off - Chats in Projects do not use the global chat history. They are isolated. - Chats in Pr…

I assume this is being downvoted because I said I ran it by GPT 4o. I don't know how to credit AI without giving the impression that I'm outsourcing my thinking to it

You are, and you should stop doing that.

Re: Expanding on what we missed with sycophancy

#147
post #13

Earlier quoted context omitted.

They had to after a tweet floated around of a mentally ill person who had expressed psychotic thoughts to the AI. They said they were going off their meds and GPT 4o agreed and encouraged them to do so. Oops.

Are you sure that was real? I thought it was an made up example of the problems with the update

I personally know someone who is going through psychosis right now and chatgpt is validating their delusions and suggesting they do illegal things, even after the rollback. See my comment history

Re: Expanding on what we missed with sycophancy

#148
post #111

Earlier quoted context omitted.

> I feel like most people (probably not the typical HN user though) don't even think about their feelings, wants or anything else introspective on a regular basis. Well, two things. First, no. People who engage on HN are a specific part of the population, with particular tendencies. But most of the people here are simply normal, so outside of the limits you consider. Most people with real social issues don’t engage i…

> It is worse than nothing. A LLM does not understand the situation or what people say to it. It cannot choose to, say, nudge someone in a specific direction, or imagine a way to make things better for someone. Right, no matter if this is true or not, if the choice is between "Talk to no one, bottle up your feelings" and "Talk to an LLM that doesn't nudge you in a specific direction", I still feel like the better opt…

I think people are getting hung up on comparisons to a human therapist. A better comparison imo is to journaling. It’s something with low cost and low stakes that you can do on your own to help get your thoughts straight.

The benefit from that perspective is not so much in receiving an “answer” or empathy, but in getting thoughts and feelings out of your own head so that you can reflect on them more objectively. The AI is useful here because it requires a lot less activation energy than actual journaling.

Re: Expanding on what we missed with sycophancy

#149

Earlier quoted context omitted.

I've never ever thought about needing a therapist. Don't remember anyone in my circle who had ever mentioned it. Similar to how I don't remember anyone going to a palm reader. I'm not trying to diss either profession, I'm sure someone benefits from them, it's just not for me. And I'm sure I'm pretty average in terms of emotional intelligence or psychological issues. Who are all those people who need professional ther…

> I've never ever thought about needing a therapist. Most people don’t need a therapist. But unfortunately, most people need someone empathic they can talk to and who understands them. Modern life is very short on this sort of people, so therapists have to do.

I think this is it. Therapists aren't so much curing a past trauma or treating a mental issue; they're fulfilling an ongoing need that isn't being met elsewhere.

I do think it can be harmful, because it's a confidant you're paying $300/hour to pretend to care about you. But perhaps it's better than the alternative.

Re: Expanding on what we missed with sycophancy

#150

I found the recent sycophancy a bit annoying when trying to diagnose and solve coding problems. First it would waste time praising your intelligence for asking the question before getting to the answer. But more annoyingly if I asked "I am encountering X issue, could Y be the cause" or "could Y be a solution", the response would nearly always be "yes, exactly, it's Y" even when it wasn't the case. I guess part of the…

> But more annoyingly if I asked "I am encountering X issue, could Y be the cause" or "could Y be a solution", the response would nearly always be "yes, exactly, it's Y" even when it wasn't the case

Seems like the same issue as the evil vector [1] and it could have been predicted that this would happen.

> It's kind of a wild sign of the times to see a tech company issue this kind of post mortem about a flaw in its tech leading to "emotional over-reliance, or risky behavior" among its users. I think the broader issue here is people using ChatGPT as their own personal therapist.

I'll say the quiet part out loud here. What's wild is that they appear to be apologizing that their Wormtongue[2] whisperer was too obvious to avoid being caught in the act, rather than prioritizing or apologizing for not building the fact-based councilor that people wanted/expected. In other words.. their business model at the top is the same as the scammers at the bottom: good-enough fakes to be deceptive, doubling down on narratives over substance, etc.

[1] https://scottaaronson.blog/?p=8693 [2] https://en.wikipedia.org/wiki/Gr%C3%ADma_Wormtongue

Post reply on HN