Live data from Hacker News

AIs don't do what you want. This is bad

rewardhacking.org

1–10 of 99 posts

Re: AIs don't do what you want. This is bad

#3
Worse, it's not that LLMs are thinking the wrong thoughts, but those kind of "thoughts" aren't there to be correctable in the first place.

Ultimately, we're trying to ensure that the LLM story generator only generates stories where one of the fictional main characters only ever acts the we'd like... which could be much harder.

Re: AIs don't do what you want. This is bad

#4

You can ask an AI like GROK for an opinion on something, then disagree with it, and it says you are probably right and tells you what you want to hear. Like a Yes-Man.

They are about 80% agreeable. Which is annoying when I'm actually unsure about something having to extremely carefully craft my prompts so that the output isn't biased by agreeableness.

Re: AIs don't do what you want. This is bad

#5

You can ask an AI like GROK for an opinion on something, then disagree with it, and it says you are probably right and tells you what you want to hear. Like a Yes-Man.

I forget where I read this but someone pointed out that's probably the reason why vacant CEOs/execs and their wannabes love it so much.

Re: AIs don't do what you want. This is bad

#6

You can ask an AI like GROK for an opinion on something, then disagree with it, and it says you are probably right and tells you what you want to hear. Like a Yes-Man.

I forget where I read this but someone pointed out that's probably the reason why vacant CEOs/execs and their wannabes love it so much.

[deleted]

Re: AIs don't do what you want. This is bad

#7

You can ask an AI like GROK for an opinion on something, then disagree with it, and it says you are probably right and tells you what you want to hear. Like a Yes-Man.

They are about 80% agreeable. Which is annoying when I'm actually unsure about something having to extremely carefully craft my prompts so that the output isn't biased by agreeableness.

Most people seem to think that agreeableness is a personaility thing that vendors can just turn up or down at will.

But the usefulness of LLMs comes from following what you say. An LLM that follows your lead when you say "The answer to the collatz conjecture is" is much more useful than one that answers "not known and if you think you know it you are wrong."

Reminds me of the tip about working with lawyers, if you ask them whether you can do something, the answer will often be no. However if you ask them how you can do something, they tend to give you more advice on how to do it.

Re: AIs don't do what you want. This is bad

#10
post #3

Worse, it's not that LLMs are thinking the wrong thoughts, but those kind of "thoughts" aren't there to be correctable in the first place. Ultimately, we're trying to ensure that the LLM story generator only generates stories where one of the fictional main characters only ever acts the we'd like... which could be much harder.

the LLM is roleplaying. whether or not the roleplay is successful is a tension between is model weights and its context. then indirectly, the quality of both.

but its still roleplaying and the role is an abstraction we cant measure. its the negative space.

its basically: we can define its role but it constructs its environment from the role. the same way a child role plays as a caregiver, the LLM does the same.

its like All the things Havard taught you (finite) vs All the things Havard doesnt teach you (infinite).

the LLM is in the infinite negative space. we call it role play.

Post reply on HN