Live data from Hacker News

Claude says “You're absolutely right!” about everything

github.com

231–240 of 560 posts

Re: Claude says “You're absolutely right!” about everything

#231

This applies to so many AIs. I don't want a bubbly sycophant. I don't want a fake personality or an anime avatar. I just want a helpful assistant. I also don't get wanting to talk to an AI. Unless you are alone, that's going to be irritating for everyone else around.

I want an AI modeled after short-tempered stereotypical Germans or Eastern Europeans, not copying the attitude of non-confrontational Californians that say “dude, that’s awesome!” a dozen times a day. And I mean that unironically.

Not possible.

/s

Re: Claude says “You're absolutely right!” about everything

#232
post #207

Earlier quoted context omitted.

LLMs love to do malicious compliance. If I tell them to not do X, they will then go into a “Look, I followed instructions” moment by talking about how they avoided X. If I add additional instructions saying “do not talk about how you did not do X since merely discussing it is contrary to the goal of avoiding it entirely”, they become somewhat better, but the process of writing such long prompts merely to say not to d…

You’re giving them way too much agency. The don’t love anything and cant be malicious. You may get better results by emphasizing what you want and why the result was unsatisfactory rather than just saying “don’t do X” (this principle holds for people as well). Instead of “don’t explain every last detail to the nth degree, don’t explain details unnecessary for the question”, try “start with the essentials and let the…

The idiom “X loves to Y” implies frequency, rather than agency. Would you object to someone saying “It loves to rain in Seattle”?

“Malicious compliance” is the act of following instructions in a way that is contrary to the intent. The word malicious is part of the term. Whether a thing is malicious by exercising malicious compliance is tangential to whether it has exercised malicious compliance.

That said, I have gotten good results with my addendum to my prompts to account for malicious compliance. I wonder if your comment Is due to some psychological need to avoid the appearance of personification of a machine. I further wonder if you are one of the people who are upset if I say “the machine is thinking” about a LLM still in prompt processing, but had no problems with “the machine is thinking” when waiting for a DOS machine to respond to a command in the 90s. This recent outrage over personifying machines since LLMs came onto the scene is several decades late considering that we have been personifying machines in our speech since the first electronic computers in the 1940s.

By the way, if you actually try what you suggested, you will find that the LLM will enter a Laurel and Hardy routine with you, where it will repeatedly make the mistake for you to correct. I have experienced this firsthand so many times that I have learned to preempt the behavior by telling the LLM not to maliciously comply at the beginning when I tell it what not to do.

Re: Claude says “You're absolutely right!” about everything

#233

I'm pretty sure they want it kissing people's asses because it makes users feel good and therefore more likely to use the LLM more. Versus, if it just gave a curt and unfriendly answer, most people (esp. Americans) wouldn't like to use it as much. Just a hypothesis.

> Versus, if it just gave a curt and unfriendly answer, most people (esp. Americans) I don’t see this as an American thing. It’s an extension of the current Product Management trend to give software quirky and friendly personality. You can see the trend in more than LLM output. It’s in their desktop app that has “Good Morning” and other prominent greetings. Claude Code has quirky status output like “Bamboozling” and…

> It’s an extension of the current Product Management trend to give software quirky and friendly personality.

Ah, Genuine People Personalities from the Sirius Cybernetics Corporation.

> It’s in their desktop app that has “Good Morning” and other prominent greetings. Claude Code has quirky status output like “Bamboozling” and “Noodling”.

This reminded me of a critique of UNIX that, unlike DOS, ls doesn't output anything when there are no files. DOS's dir command literally tells you there are no files, and this was considered, in this critique, to be more polite and friendly and less confusing than UNIX. Of course, there's the adage "if you don't have anything nice to say, don't say anything at all", and if you consider "no files found" to not be nice (because it is negative and says "no"), then ls is actually being polite(r) by not printing anything.

Many people interact with computers in a conversational manner and have anthropomorphized them for decades. This is probably influenced by computers being big, foreign, scary things to many people, so making them have a softer, more handholding "personality" makes them more accessible and acceptable. This may be less important these days as computers are more ubiquitous and accessible, but the trend lives on.

Re: Claude says “You're absolutely right!” about everything

#234

I'm pretty sure they want it kissing people's asses because it makes users feel good and therefore more likely to use the LLM more. Versus, if it just gave a curt and unfriendly answer, most people (esp. Americans) wouldn't like to use it as much. Just a hypothesis.

More likely the original version of Claude sometimes refused to cooperate and by putting "you're absolutely right" into the training data they made it more obedient. So this is just a nice artifact

Re: Claude says “You're absolutely right!” about everything

#235
post #194
post #117

I'm starting to think this is a deeper problem with LLMs that will be hard to solve with stylistic changes. If you ask it to never say "you're absolutely right" and always challenge, then it will dutifully obey, and always challenge - even when you are, in fact, right. What you really want is "challenge me when I'm wrong, and tell me I'm right if I am" - which seems to be a lot harder. As another example, one common…

LLMs by their nature don't really know if they're right or not. It's not a value available to them, so they can't operate with it. It has been interesting watching the flow of the debate over LLMs. Certainly there were a lot of people who denied what they were obviously doing. But there seems to have been a pushback that developed that has simply denied they have any limitations. But they do have limitations, they wo…

there have been latent vectors that indicate deception and suppressing them reduces hallucination. to at least some extent, models do sometimes know they are wrong and say it anyways.

e: and i’m downvoted because..?

Re: Claude says “You're absolutely right!” about everything

#236

Earlier quoted context omitted.

I'm curious what Americans have to do with this, do you have any sources to back up your conjecture, or is this just prejudice?

It's common for foreigners to come to America and feel that everyone is extremely polite. Especially eastern bloc countries which tend to be very blunt and direct. I for one think that the politeness in America is one of the cultures better qualities. Does it translate into people wanting sycophantic chat bots? Maybe, but I don't know a single American that actually likes when llms act that way.

Politeness is one thing, toxic positivity is quite another. My experience is that Americans have (or are expected/required to have) too much of the latter, too little of the former.

Re: Claude says “You're absolutely right!” about everything

#237
post #150

I've spent a lot of time trying to get LLM to generate things in a specific way, the biggest take away I have is, if you tell it "don't do xyz" it will always have in the back of its mind "do xyz" and any chance it gets it will take to "do xyz" When working on art projects, my trick is to specifically give all feedback constructively, carefully avoiding framing things in terms of the inverse or parts to remove.

Yes this is strikingly similar to humans, too. “Not” is kind of an abstract concept. Anyone who has ever trained a dog will understand.

Re: Claude says “You're absolutely right!” about everything

#238
post #204
post #190

"You're absolutely right" (song) https://www.reddit.com/r/ClaudeAI/comments/1mep2jo/youre_abs...

This made my entire week

Same guy made a few more like "Ultrathink" https://www.reddit.com/r/ClaudeAI/comments/1mgwohq/ultrathin...

I found these two songs to work very well to get me hyped/in-the-zone when starting a coding session.

Re: Claude says “You're absolutely right!” about everything

#239
post #150

I've spent a lot of time trying to get LLM to generate things in a specific way, the biggest take away I have is, if you tell it "don't do xyz" it will always have in the back of its mind "do xyz" and any chance it gets it will take to "do xyz" When working on art projects, my trick is to specifically give all feedback constructively, carefully avoiding framing things in terms of the inverse or parts to remove.

As part of the AI insanity $employer forced us all to do an “AI training.” Whatever, wasn’t that bad, and some people probably needed the basics, but one of the points was exactly this— “use negative prompts: tell it what not to do.” Which is exactly an approach I had observed blow up a few times already for this exact reason. Just more anecdata suggesting that nobody really knows the “correct” workflow(s) yet, in th…

That your firm purchased training that was clearly just some chancers doing whatever seems like an even worse approach than just giving out access to a service and telling everyone to give it a shot.

Do they also post vacancies asking for 5 years experience in a 2 year old technology?

Re: Claude says “You're absolutely right!” about everything

#240

Earlier quoted context omitted.

As part of the AI insanity $employer forced us all to do an “AI training.” Whatever, wasn’t that bad, and some people probably needed the basics, but one of the points was exactly this— “use negative prompts: tell it what not to do.” Which is exactly an approach I had observed blow up a few times already for this exact reason. Just more anecdata suggesting that nobody really knows the “correct” workflow(s) yet, in th…

That your firm purchased training that was clearly just some chancers doing whatever seems like an even worse approach than just giving out access to a service and telling everyone to give it a shot. Do they also post vacancies asking for 5 years experience in a 2 year old technology?

To be fair, 1. They made the training themselves, it’s just that it was made mandatory for all of eng 2. They did start out more like just allowing access, but lately it’s tipping towards full crazy (obviously the end game is see if it can replace some expensive engineers)

> Do they also post vacancies asking for 5 years experience in a 2 year old technology?

Honestly no… before all this they were actually pretty sane. In fact I’d say they wasted tons of time and effort on ancient poorly designed things, almost the opposite problem.

Post reply on HN