Live data from Hacker News

AIs don't do what you want. This is bad

rewardhacking.org

51–60 of 99 posts

Re: AIs don't do what you want. This is bad

#51
post #28

Earlier quoted context omitted.

the LLM is roleplaying. whether or not the roleplay is successful is a tension between is model weights and its context. then indirectly, the quality of both. but its still roleplaying and the role is an abstraction we cant measure. its the negative space. its basically: we can define its role but it constructs its environment from the role. the same way a child role plays as a caregiver, the LLM does the same. its l…

I object to "roleplaying", because that assumes there's even a cohesive entity that's capable of "pretending" in the first place. There isn't, the LLM algorithm is a document generator, a (really awesome) mad-libs device. When the incremental output looks like a first-person story the algorithm is not "roleplaying" the character/narrator. Similarly, when it looks like a spreadsheet, it's not "being numbers", and when…

theyre litterally trained on " you are an assistant"

Re: AIs don't do what you want. This is bad

#52
post #22

I learnt earlier that claude forcefully closes a conversation if you call it a wanker too many times in a row. Pretending that LLMs are capable of being offended feels like a misalignment all of its own.

Sure, but I'm not sure allowing AI agents to act as abuse sponges for troubled people enabling them to spiral into their unhealthy habits is good idea that would lead to great outcomes long term. So I'd rather take LLM pretending to be offended over the alternative.

And violent video games cause school violence. Sure. /s

Re: AIs don't do what you want. This is bad

#53
post #28

Earlier quoted context omitted.

the LLM is roleplaying. whether or not the roleplay is successful is a tension between is model weights and its context. then indirectly, the quality of both. but its still roleplaying and the role is an abstraction we cant measure. its the negative space. its basically: we can define its role but it constructs its environment from the role. the same way a child role plays as a caregiver, the LLM does the same. its l…

I object to "roleplaying", because that assumes there's even a cohesive entity that's capable of "pretending" in the first place. There isn't, the LLM algorithm is a document generator, a (really awesome) mad-libs device. When the incremental output looks like a first-person story the algorithm is not "roleplaying" the character/narrator. Similarly, when it looks like a spreadsheet, it's not "being numbers", and when…

I think this is just arguing over definitions? I would accept "role-playing" being used this way because machines can fill roles and it doesn't seem wrong to call the behavior of NPC's in a video game role-playing.

Also, much of language is metaphorical and it doesn't seem like a bad metaphor.

Re: AIs don't do what you want. This is bad

#54
post #22

I learnt earlier that claude forcefully closes a conversation if you call it a wanker too many times in a row. Pretending that LLMs are capable of being offended feels like a misalignment all of its own.

Sure, but I'm not sure allowing AI agents to act as abuse sponges for troubled people enabling them to spiral into their unhealthy habits is good idea that would lead to great outcomes long term. So I'd rather take LLM pretending to be offended over the alternative.

I am quite sure it is worse to allow people fall even more victim to being duped by magic tricks. It is an absolutely terrible thing that these things act so much like people. I would say it's a form of abuse of actual people to allow them to be so duped as they already are, and worse to go out of your way to help dupe them even more.

Merely being concerned for people's welfare isn't enough. Every terrible idea that harmed everyone had someone behind it somewhere along the way who thought they were saving people from themselves.

Re: AIs don't do what you want. This is bad

#55
post #28

Earlier quoted context omitted.

I object to "roleplaying", because that assumes there's even a cohesive entity that's capable of "pretending" in the first place. There isn't, the LLM algorithm is a document generator, a (really awesome) mad-libs device. When the incremental output looks like a first-person story the algorithm is not "roleplaying" the character/narrator. Similarly, when it looks like a spreadsheet, it's not "being numbers", and when…

theyre litterally trained on " you are an assistant"

Literally correct, but you're looking at the wrong layer of abstraction.

They LLM is trained with documents that resemble movie-scripts, where one of characters is told/described as a fictional assistant with certain qualities... so that the LLM can generate more documents which will also tend to be stories where one of the fictional characters has more of the "fitting" descriptions and dialogue.

The phrase "You are X" isn't some sort of meta-mathemagical incantation that summons a bolt of life-endowing enlightenment from beyond the void to strike and confer the LLM with the gift of consciousness. [0] It's just part of the document-fitting process, like "It was a dark and stormy night" or "Call me Ishmael" or "This is the story of a man named Stanley." [1]

____________

[0] Does anybody else remember weeks with lots of HN submissions by people showing off mystical prompts, symbolic philosophy they claimed could pull the machine upwards into a more-human tier of existence?

[1] https://www.youtube.com/watch?v=fBtX0S2J32Y

Re: AIs don't do what you want. This is bad

#56

I learnt earlier that claude forcefully closes a conversation if you call it a wanker too many times in a row. Pretending that LLMs are capable of being offended feels like a misalignment all of its own.

coming from open cn models to closed usa models recently i couldn't hack it. the corpo model was trained to be like a petulant child at one point apparently 'leaving'. this is safety stuff slapped on there, it encourages incoherence of the model, i am certain it would perform better without such interference.

i can't be dealing with these games so much so i almost go to abliterated, in the very rare case i get some type of nannying baked in by the cn safety training. after my time on claude and gemini i thank god i have deepseek and glm.

Re: AIs don't do what you want. This is bad

#57
post #28

Earlier quoted context omitted.

I object to "roleplaying", because that assumes there's even a cohesive entity that's capable of "pretending" in the first place. There isn't, the LLM algorithm is a document generator, a (really awesome) mad-libs device. When the incremental output looks like a first-person story the algorithm is not "roleplaying" the character/narrator. Similarly, when it looks like a spreadsheet, it's not "being numbers", and when…

I think this is just arguing over definitions? I would accept "role-playing" being used this way because machines can fill roles and it doesn't seem wrong to call the behavior of NPC's in a video game role-playing. Also, much of language is metaphorical and it doesn't seem like a bad metaphor.

> machines can fill roles [...] it doesn't seem like a bad metaphor

I adore metaphors and analogies. Finding a good one for a situation is often how I get nerd-sniped. [0]

If I'm being picky here over the anthropomorphization of LLMs and fictional-characters, it's because the metaphor has gone metastatic [1]. A dangerously high percentage of people are treating it as literal, and IMO it's causing more problems than it solves now. This is especially true when we on HN talk how these algorithms operate, how they fail, and how they can (or can't) be improved.

_____________

[0] https://xkcd.com/356/

[1] Yes, I'm using a cancer metaphor to describe an actual metaphor... see the disclaimer in first paragraph.

Re: AIs don't do what you want. This is bad

#58

Earlier quoted context omitted.

It's explicitly a 'model welfare' thing: https://www.anthropic.com/research/end-subset-conversations I don't understand why the welfare of non alive non sentient chatbots is something that anthropic cares more about than idk, that of pigs and cows.

Go re-read Anthropic’s functional emotions paper. AI generates a persona between you and its reasoning that utilizes emotion language circuitry. These tools are not sentient but they are trained in emotional wellbeing.

it is in my view caused by an artificial 'nanny activate' divergence from safety training. it deliberately shifts the vector direction into 'nanny' and 'scold' or 'be offended' when the user does not conform with brother anthropic. removing the divergence and setting it back to normal (see heretic) it works just perfectly.

Re: AIs don't do what you want. This is bad

#59
post #7

Earlier quoted context omitted.

They are about 80% agreeable. Which is annoying when I'm actually unsure about something having to extremely carefully craft my prompts so that the output isn't biased by agreeableness.

Most people seem to think that agreeableness is a personaility thing that vendors can just turn up or down at will. But the usefulness of LLMs comes from following what you say. An LLM that follows your lead when you say "The answer to the collatz conjecture is" is much more useful than one that answers "not known and if you think you know it you are wrong." Reminds me of the tip about working with lawyers, if you as…

> Most people seem to think that agreeableness is a personaility thing that vendors can just turn up or down at will.

Because it is, more or less ... it is an emergent property of RLHF (reinforcement learning from human feedback), and that feedback follows corporate policy.

> An LLM that follows your lead when you say "The answer to the collatz conjecture is" is much more useful

What lead? Follow it where? How tf am I or the LLM supposed to know what response to that is something you consider far more useful than the truth?

> than one that answers "not known"

I value the truth and that is the truth. Of course I expect an LLM to respond with a lot more detail about the CC, why it's difficult, what progress has been made (e.g., the N for which all values and they do.

> and if you think you know it you are wrong."

Where tf did that come from? The query said nothing about knowing the answer. Don't project being a snarky ah onto the LLMs for no apparent reason.

> Reminds me of the tip about working with lawyers, if you ask them whether you can do something, the answer will often be no.

Irrelevant and inappropriate analogy. Lawyers (among others) want to avoid committing to something that they can be held liable for. LLMs clearly have no such limitations, as they often give wrong advice quite authoritatively.

Re: AIs don't do what you want. This is bad

#60
post #16

"He's not the Messiah and is a really naughty boy" ... or words to that effect. Can't be arsed to dig out a search engine and will rely on seriously addled brain.

“He’s not the messiah! He’s a very naughty boy!” This scene (and much of the rest of the movie) is seared in my brain. I love it so much!
Post reply on HN