Live data from Hacker News

Claude says “You're absolutely right!” about everything

github.com

191–200 of 560 posts

Re: Claude says “You're absolutely right!” about everything

#191
As a neurodiverse British person I tend to communicate more directly than the average English speaker and I find LLM's manner of speech very off-putting and insincere, which in some cases it literally is. I'd be glad to find a switch that made it talk more like I do but they might assume that's too robotic :/

Re: Claude says “You're absolutely right!” about everything

#192
post #117

I'm starting to think this is a deeper problem with LLMs that will be hard to solve with stylistic changes. If you ask it to never say "you're absolutely right" and always challenge, then it will dutifully obey, and always challenge - even when you are, in fact, right. What you really want is "challenge me when I'm wrong, and tell me I'm right if I am" - which seems to be a lot harder. As another example, one common…

I'm more reminded of Tom Scott's talk at the Royal Institution "There is no Algorithm for Truth"[0]. A lot of what you're talking about is the ability to detect Truth, or even truth! [0] https://www.youtube.com/watch?v=leX541Dr2rU

> I'm more reminded of Tom Scott's talk at the Royal Institution "There is no Algorithm for Truth"[0].

Isn't there?

https://en.wikipedia.org/wiki/Solomonoff%27s_theory_of_induc...

Re: Claude says “You're absolutely right!” about everything

#193

I'm pretty sure they want it kissing people's asses because it makes users feel good and therefore more likely to use the LLM more. Versus, if it just gave a curt and unfriendly answer, most people (esp. Americans) wouldn't like to use it as much. Just a hypothesis.

I'm curious what Americans have to do with this, do you have any sources to back up your conjecture, or is this just prejudice?

Prejudice, based on my anecdotal experience. I live in the US but have spent a decent amount of time in Europe (mostly Germany).

Re: Claude says “You're absolutely right!” about everything

#194
post #117

I'm starting to think this is a deeper problem with LLMs that will be hard to solve with stylistic changes. If you ask it to never say "you're absolutely right" and always challenge, then it will dutifully obey, and always challenge - even when you are, in fact, right. What you really want is "challenge me when I'm wrong, and tell me I'm right if I am" - which seems to be a lot harder. As another example, one common…

LLMs by their nature don't really know if they're right or not. It's not a value available to them, so they can't operate with it.

It has been interesting watching the flow of the debate over LLMs. Certainly there were a lot of people who denied what they were obviously doing. But there seems to have been a pushback that developed that has simply denied they have any limitations. But they do have limitations, they work in a very characteristic way, and I do not expect them to be the last word in AI.

And this is one of the limitations. They don't really know if they're right. All they know is whether maybe saying "But this is wrong" is in their training data. But it's still just some words that seem to fit this situation.

This is, if you like and if it helps to think about it, not their "fault". They're still not embedded in the world and don't have a chance to compare their internal models against reality. Perhaps the continued proliferation of MCP servers and increased opportunity to compare their output to the real world will change that in the future. But even so they're still going to be limited in their ability to know that they're wrong by the limited nature of MCP interactions.

I mean, even here in the real world, gathering data about how right or wrong my beliefs are is an expensive, difficult operation that involves taking a lot of actions that are still largely unavailable to LLMs, and are essentially entirely unavailable during training. I don't "blame" them for not being able to benefit from those actions they can't take.

Re: Claude says “You're absolutely right!” about everything

#195
post #150

I've spent a lot of time trying to get LLM to generate things in a specific way, the biggest take away I have is, if you tell it "don't do xyz" it will always have in the back of its mind "do xyz" and any chance it gets it will take to "do xyz" When working on art projects, my trick is to specifically give all feedback constructively, carefully avoiding framing things in terms of the inverse or parts to remove.

This is a childrearing technique, too: say “please do X”, where X precludes Y, rather than saying “please don’t do Y!”, which just increases the salience, and therefore likelihood, of Y.

Re: Claude says “You're absolutely right!” about everything

#196
post #183

Earlier quoted context omitted.

I have this same problem. I’ve added a bunch of instructuons to try and stop ChatGPT being so sycophantic, and now it always mentions something about how it’s going to be ‘straight to the point’ or give me a ‘no bs version’. So now I just have that as the intro instead of ‘that’s a sharp observation’

> it always mentions something about how it’s going to be ‘straight to the point’ or give me a ‘no bs version’ That's how you suck up to somebody who doesn't want to see themselves as somebody you can suck up to. How does an LLM know how to be sycophantic to somebody who doesn't (think they) like sycophants? Whether it's a naturally emergent phenomenon in LLMs or specifically a result of its corporate environment, I'…

Garbage in, garbage out.

It's that simple.

Re: Claude says “You're absolutely right!” about everything

#198
post #150

I've spent a lot of time trying to get LLM to generate things in a specific way, the biggest take away I have is, if you tell it "don't do xyz" it will always have in the back of its mind "do xyz" and any chance it gets it will take to "do xyz" When working on art projects, my trick is to specifically give all feedback constructively, carefully avoiding framing things in terms of the inverse or parts to remove.

> the biggest take away I have is, if you tell it "don't do xyz" it will always have in the back of its mind "do xyz" and any chance it gets it will take to "do xyz" You're absolutely right! This can actually extend even to things like safety guardrails. If you tell or even train an AI to not be Mecha-Hitler, you're indirectly raising the probability that it might sometimes go Mecha-Hitler. It's one of many reasons w…

> You're absolutely right!

Is this irony, actual LLM output or another example of humans adopting LLM communication patterns?

Re: Claude says “You're absolutely right!” about everything

#199

I'm pretty sure they want it kissing people's asses because it makes users feel good and therefore more likely to use the LLM more. Versus, if it just gave a curt and unfriendly answer, most people (esp. Americans) wouldn't like to use it as much. Just a hypothesis.

As a Finn, it makes me want to use it much, much less if it kisses ass.

Finns need to mentally evolve beyond this mindset.

Somebody being polite and friendly to you does not mean that the person is inferior to you and that you should therefore despise them.

Likewise somebody being rude and domineering to you does not mean that they are superior to you and should be obeyed and respected.

Politeness is a tool and a lubricant, and Finns probably loose out on a lot of international business and opportunities because of this mentality that you're demonstrating. Look at the Japanese for inspiration, who were an economic miracle, while sharing many positive values with the Finns.

Re: Claude says “You're absolutely right!” about everything

#200

Earlier quoted context omitted.

I have this same problem. I’ve added a bunch of instructuons to try and stop ChatGPT being so sycophantic, and now it always mentions something about how it’s going to be ‘straight to the point’ or give me a ‘no bs version’. So now I just have that as the intro instead of ‘that’s a sharp observation’

Any time you're fighting the training + system prompt with your own instructions and prompting the results are going to be poor, and both of those things are heavily geared towards being a cheery and chatty assistant.

[dead]
Post reply on HN