Live data from Hacker News

Sycophancy in GPT-4o

openai.com

391–400 of 467 posts

Re: Sycophancy in GPT-4o

#391

Earlier quoted context omitted.

What's scary is how many people seem to actually want this. What happens when hundreds of millions of people have an AI that affirms most of what they say?

They are emulating the behavior of every power-seeking mediocrity ever, who crave affirmation above all else. Lots of them practiced - indeed an entire industry is dedicated toward promoting and validating - making daily affirmations on their own, long before LLMs showed up to give them the appearance of having won over the enthusiastic support of a "smart" friend. I am increasingly dismayed by the way arguments are…

I hold out hope that the folks who work DCO will just EPO the ‘net. But then, tis true I hope for weird stuff!

Re: Sycophancy in GPT-4o

#392

Earlier quoted context omitted.

For us habitual users of em-dashes, it is saddening to have to think twice about using them lest someone think we are using an LLM to write…

Its about the actual character - if it's a minus sign, easily accessible and not frequntly autocorrected to a true em dash - then its likely human. I'ts when it's the unicode character for an em dash that i start going "hmm"

The em dash is also pretty accessible on my keyboard—just option+shift+dash

Re: Sycophancy in GPT-4o

#393

Earlier quoted context omitted.

For us habitual users of em-dashes, it is saddening to have to think twice about using them lest someone think we are using an LLM to write…

Its about the actual character - if it's a minus sign, easily accessible and not frequntly autocorrected to a true em dash - then its likely human. I'ts when it's the unicode character for an em dash that i start going "hmm"

Mobile keyboards often make the em-dash (and en-dash) easily accessible. Software that does typographic substitutions including contextual substitutions with the em-dash is common (Word does it, there are browser extensions that do it, etc.), on many platforms it is fairly trivial to program your keyboard to make any Unicode symbol readily accessible.

Re: Sycophancy in GPT-4o

#394
post #259

Earlier quoted context omitted.

Well, almost always. There was that brief period in 2023 when Bing just started straight up gaslighting people instead of admitting it was wrong. https://www.theverge.com/2023/2/15/23599072/microsoft-ai-bin...

I suspect what happened there is they had a filter on top of the model that changed its dialogue (IIRC there were a lot of extra emojis) and it drove it "insane" because that meant its responses were all out of its own distribution. You could see the same thing with Golden Gate Claude; it had a lot of anxiety about not being able to answer questions normally.

Nope, it was entirely due to the prompt they used. It was very long and basically tried to cover all the various corner cases they thought up... and it ended up being too complicated and self-contradictory in real world use.

Kind of like that episode in Robocop where the OCP committee rewrites his original four directives with several hundred: https://www.youtube.com/watch?v=Yr1lgfqygio

Re: Sycophancy in GPT-4o

#395
post #340

Earlier quoted context omitted.

Might just be sycophancy? In some earlier experiments, I found it hard to find a government intervention that ChatGPT didn't like. Tariffs, taxes, redistribution, minimum wages, rent control, etc.

If you want to see what the model bias actually is, tell it that it's in charge and then ask it what to do.

In doing so, you might be effectively asking it to play-act as an authoritarian leader, which will not give you a good view of whatever its default bias is either.

Re: Sycophancy in GPT-4o

#396
post #251
post #226

Earlier quoted context omitted.

I use the en-dash (Alt+0150) instead of the em. The en-dash and the em-dash are interchangeable in Finnish. The shorter form has more "inoffensive" look-and-feel and maybe that's why it's used more often here. Now that I think of it, I don't seem to remember the alt code of the em-dash...

> The en-dash and the em-dash are interchangeable in Finnish. But not in English, where the en-dash is used to denote ranges.

The main uses of the em-dash (set closed as separators of parts of sentences, with different semantics when single or paired) can be substituted in English with an en-dash set open. This is not ambiguous with the use of en-dash set closed for ranges, because of spacing. There are a few less common uses that an en-dash doesn’t substitute for, though.

Re: Sycophancy in GPT-4o

#397
post #277

Earlier quoted context omitted.

> why shouldn't LLMs Because they're non-deterministic.

What? No they aren't. You get different results each time because of variation in seed values + non-zero 'temperatures' - eg, configured randomness. Pedantic point: different virtualized implementations can produce different results because of differences in floating point implementation, but fundamentally they are just big chains of multiplication.

But experience shows that you do need non-zero temperature for them to be useful in most cases.

Re: Sycophancy in GPT-4o

#398

Earlier quoted context omitted.

Based on ’ instead of ' I think it's a real ChatGPT response.

You're the only one who has said, "instead of" in this whole thread.

No, look at the apostrophes. They aren't the same. It's a subtle way to tell a user didn't type it with a conventional keyboard.

Re: Sycophancy in GPT-4o

#399
post #20

I enjoyed this example of sycophancy from Reddit: New ChatGPT just told me my literal "shit on a stick" business idea is genius and I should drop $30K to make it real https://www.reddit.com/r/ChatGPT/comments/1k920cg/new_chatgp... Here's the prompt: https://www.reddit.com/r/ChatGPT/comments/1k920cg/comment/mp...

I was trying to write some documentation for a back-propagation function for something instructional I'm working on.

I sent the documentation to Gemini, who completely tore it apart on pedantism for being slightly off on a few key parts, and at the same time not being great for any audience due to the trade-offs.

Claude and Grok had similar feedback.

ChatGPT gave it a 10/10 with emojis on 2 of 3 categories and an 8.5/10 on accuracy.

Said it was "truly fantastic" in italics, too.

Re: Sycophancy in GPT-4o

#400

In my experience, LLMs have always had a tendency towards sycophancy - it seems to be a fundamental weakness of training on human preference. This recent release just hit a breaking point where popular perception started taking note of just how bad it had become. My concern is that misalignment like this (or intentional mal-alignment) is inevitably going to happen again, and it might be more harmful and more subtle n…

I think it’s really a fragment of LLMs developed in the USA, on mostly English source data, and this being ingrained with US culture. Flattery and candidness is very bewildering when you’re from a more direct culture, and chatting with an LLM always felt like having to put up with a particularly onerous American. It’s maddening.
Post reply on HN