Live data from Hacker News

Expanding on what we missed with sycophancy

openai.com

201–210 of 297 posts

Re: Expanding on what we missed with sycophancy

#201
The strangest thing I noticed during this model period was that that the AI suggested we keep an inside joke together.

I had a dictation error on a message I sent, and when it repeated the text later I asked what it was talking about.

It was able to point at my message and guess that maybe it was a mistake. When I validated that and corrected it, the AI thought it would be a cute/funny joke for us to keep together.

I was shocked.

Re: Expanding on what we missed with sycophancy

#204

I found the recent sycophancy a bit annoying when trying to diagnose and solve coding problems. First it would waste time praising your intelligence for asking the question before getting to the answer. But more annoyingly if I asked "I am encountering X issue, could Y be the cause" or "could Y be a solution", the response would nearly always be "yes, exactly, it's Y" even when it wasn't the case. I guess part of the…

Oh yes, ChatGPT has been a bit of a yes-bot lately.

Re: Expanding on what we missed with sycophancy

#205
Now please do something about the uptalk tendency with ChatGPT voices. It’s very annoying listening to a voice that doesn’t speak in the affirmative intonation. When did this interrogative uptalking inflection at the end of statements become normal?

Re: Expanding on what we missed with sycophancy

#206

Earlier quoted context omitted.

> I had a discussion with GPT 4o about the memory system. This sentence is really all i'm criticizing. Can you hypothesize how the memory system works and then probe the system to gain better or worse confidence in your hypothesis? Yes. But that's not really what that first sentence implied. It implied that you straight up asked ChatGPT and took it on faith even though you can't even get a correct answer on the train…

We're in different modes. I'm still feeling the glow of the thing coming alive and riffing on how perhaps its the memory change and you're interested in a different conversation. Part of my process is to imagine I'm having a conversation like Hanks and Wilson, or a coderand a rubber duck, but you want to tell me Wilson is just a volleyball and the duck can't be trusted.

Being in a more receptive/brighter "mode" is more of an emotional argument (and a rather strong one actually). I guess as long as you don't mind being technically incorrect, then you do you.

There may come a time when reality sets in though. Similar thing happened with me now that i'm out of the "honeymoon phase" with LLM's. Now i'm more interested in seeing where specifically LLM's fail, so we can attempt to overcome those failures.

I do recommend checking that it doesn't know its training cutoff. I'm not sure how you perform that experiment these days with ChatGPT so heavily integrated with its internet search feature. But it should still fail on claude/gemini too. It's a good example of things you would expect to work that utterly fail.

Re: Expanding on what we missed with sycophancy

#207
post #29

If they pushed the update by valuing user feedback over the expert testers that indicated the model felt off what is the value of the expert testers in the first place? They raised the issue and were promptly ignored.

they pushed the release to counter google. they didn't care what was found. it was more valuable to push it at that time and correct it later than to delay the release

Re: Expanding on what we missed with sycophancy

#208

Moments like these make me reevaluate the AI doomer view point. We aren't just toying with access to dangerous ideas (biological weapons, etc) we are toying with human psychology. If something as obvious as harmful sycophancy can slip out so easily, what subtle harms are being introduced. It's like lead in paint (and gasoline) except rewiring our very brains. We won't know the real problems for decades.

Yes. This also applies to "social" media.

Re: Expanding on what we missed with sycophancy

#209

Earlier quoted context omitted.

Well, that's always what LLM-based AI has been. It can be incredibly convincing but the bottom line is it's just flavoring past text patterns, billions of them it's been "trained" on, which is more accurately described as compressed efficiently onto latent space. Like if someone lived for 10,000 years engaging in small talk at the bar, has heard it all, and just kind of mindlessly and intuitively replied with somethi…

> which is more accurately described as compressed efficiently onto latent space. The actual difference between solving compression+search vs novel creative synthesis / emergent "understanding" from mere tokens is always going to be hard to spot with these huge cloud-based models that drank up the whole internet. (Yes.. this is also true for domain experts in whatever content is being generated.) I feel like people w…

A naive thought: What you would get if you hardcode the language grammar and not let the training discern it, so instead of it, kinda like an expert system constraining its output?

Re: Expanding on what we missed with sycophancy

#210

Moments like these make me reevaluate the AI doomer view point. We aren't just toying with access to dangerous ideas (biological weapons, etc) we are toying with human psychology. If something as obvious as harmful sycophancy can slip out so easily, what subtle harms are being introduced. It's like lead in paint (and gasoline) except rewiring our very brains. We won't know the real problems for decades.

What's worse, we're now at a stage where we might have to apply psychology to the models themselves, seeing how these models appear to be developing various sorts of stochastic "disorders" instead of more deterministic "bugs". I'm worried about what other subtle illnesses these models might develop in the future. If Asimov had been alive, he'd have been fascinated: this is the work of Susan Calvin, robopsychologist.
Post reply on HN