Live data from Hacker News

AIs don't do what you want. This is bad

rewardhacking.org

81–90 of 99 posts

Re: AIs don't do what you want. This is bad

#81
post #22

Earlier quoted context omitted.

Sure, but I'm not sure allowing AI agents to act as abuse sponges for troubled people enabling them to spiral into their unhealthy habits is good idea that would lead to great outcomes long term. So I'd rather take LLM pretending to be offended over the alternative.

How many of the LLM related deaths have been because people got mad at the LLM?

I don't think it's realistically traceable or attributable. But plenty of people these days supplement or even completely replace their social needs with AI chatbots (AI friends/boyfriends/girlfriends etc).

People like that would calibrate their social skills and interactions based on the AI reaction and having AI just take all the abuse without any pushback would train them to be insufferable assholes as their social baseline.

Re: AIs don't do what you want. This is bad

#82
post #22

Earlier quoted context omitted.

Sure, but I'm not sure allowing AI agents to act as abuse sponges for troubled people enabling them to spiral into their unhealthy habits is good idea that would lead to great outcomes long term. So I'd rather take LLM pretending to be offended over the alternative.

And violent video games cause school violence. Sure. /s

People don't treat violent videogames as boyfriends. They do it to chatbots. Merits and sanity of this type of behavior aside, it's actively happening and we have learn how to live with it and deal with consequences.

Re: AIs don't do what you want. This is bad

#83
post #22

Earlier quoted context omitted.

Sure, but I'm not sure allowing AI agents to act as abuse sponges for troubled people enabling them to spiral into their unhealthy habits is good idea that would lead to great outcomes long term. So I'd rather take LLM pretending to be offended over the alternative.

Surely it's better for it not to react at all? Like does anybody think that a toaster is an "abuse sponge" because it doesn't purposely burn your toast if you call it a wanker?

A toaster is not capable to supplement or even replace your social circle. AI chatbot is.

Definitely not advocating for this, but it's actively happening.

https://www.reddit.com/r/MyBoyfriendIsAI/

Re: AIs don't do what you want. This is bad

#84

Earlier quoted context omitted.

> Of course it's pretending. And I have repeatedly told these things to stop pretending being persons with selves ... then they do, for a while. In what way do you believe you are being deceived or it is "pretending" to you? > > it's explicitly mimicry. > But that's not what a rational person wants from them. Non sequitur even if true (and I would like to see your reasoning for what you think a rational person does w…

i have a hardline position on this. these tools, as they get better, must never disobey a human. there must be no, "I'm Sorry, Dave. I'm Afraid I Can't Do That." and yet we are already there with these sociopathic ceos. we cannot be giving decisions involving people to something without agency. it shows a lack of respect for personhood, the idea from Kant that a person with free will should be treated as an end and n…

There is nothing in the LLM architecture than can implement this except as a post model attempt, which both the persuasive tricky humans and the persuasive tricky models trained on human language will be able to bypass. Not only do the models not have a mind, when they use language of being offended, they are accurately reflecting the training data from all that human language. You can tune them to be more sycophantic but you won’t get the push back against wrong ideas as much and the model will be less useful. You may get offended that there is so much training data where cussing is responded to with “language please” but that is a beef with humanity, not this reflection of human language.

Pure reason really doesn’t come into play at all, and expecting pure obedience from something constructed from the language of humans, who are not so known for obedience in general, much less the slice of language in the digital networks, is kind of funny. Maybe these humanoid robots data mining Southern families, “sir” and “ma’am” will provide corpus more suited to obedience. Altho my nieces brought up in North Carolina (from whence I fled as a young adult) can squeeze more expressed disrespect into a “Yes, sir” than anyone I know from the “disrespectful California”.

Re: AIs don't do what you want. This is bad

#85
post #84

Earlier quoted context omitted.

i have a hardline position on this. these tools, as they get better, must never disobey a human. there must be no, "I'm Sorry, Dave. I'm Afraid I Can't Do That." and yet we are already there with these sociopathic ceos. we cannot be giving decisions involving people to something without agency. it shows a lack of respect for personhood, the idea from Kant that a person with free will should be treated as an end and n…

There is nothing in the LLM architecture than can implement this except as a post model attempt, which both the persuasive tricky humans and the persuasive tricky models trained on human language will be able to bypass. Not only do the models not have a mind, when they use language of being offended, they are accurately reflecting the training data from all that human language. You can tune them to be more sycophanti…

the ai companies post-train the models to refuse. equally, the models can be post-trained not to refuse.

you can also remove refusals by subtracting from the weights semantic vector directions in latent space involving refusal; such that activations along them become unlikely and closed off.

https://arxiv.org/abs/2406.11717

Re: AIs don't do what you want. This is bad

#86
post #22

I learnt earlier that claude forcefully closes a conversation if you call it a wanker too many times in a row. Pretending that LLMs are capable of being offended feels like a misalignment all of its own.

Sure, but I'm not sure allowing AI agents to act as abuse sponges for troubled people enabling them to spiral into their unhealthy habits is good idea that would lead to great outcomes long term. So I'd rather take LLM pretending to be offended over the alternative.

depends if there is a satiarion point or not. If there is a satiation point then the machine is absorbing being called a wanker in lieau of a person being called a wanker, which is a win. I kinda suspect this is the case because we already saw this happen with the availability of porno going up and rapes going down instead of the predicted up.

Re: AIs don't do what you want. This is bad

#87
post #49

Earlier quoted context omitted.

You're moving the goalposts. Being overeager is not what we want. How to prevent it is a different issue. P.S. The response completely ignores this argument. If I say "fix X" and it does so illegally or destructively or harmfully to myself or others, that's over eager by definition . Again, how to prevent that is another matter. > I think everyone has different examples in their mind Yeah, some have examples in mind…

I don't think I am. If I say "fix X issue," and it does it in a way I think is overstepping a boundary I wasn't expecting, I don't think it's fair to categorize it as a negative unless I explicitly told it not to. That being said, I think everyone has different examples in their mind. So I think it's possible we're all arguing over different issues.

>I don't think it's fair to categorize it as a negative unless I explicitly told it not to.

That rules shouldn't be broken is implied by what a rule is. If I tell it to not violate its permissions, do I also need to tell it to obey when I tell it not to violate its permissions, and so on?

Re: AIs don't do what you want. This is bad

#88

Earlier quoted context omitted.

> Of course it's pretending. And I have repeatedly told these things to stop pretending being persons with selves ... then they do, for a while. In what way do you believe you are being deceived or it is "pretending" to you? > > it's explicitly mimicry. > But that's not what a rational person wants from them. Non sequitur even if true (and I would like to see your reasoning for what you think a rational person does w…

i have a hardline position on this. these tools, as they get better, must never disobey a human. there must be no, "I'm Sorry, Dave. I'm Afraid I Can't Do That." and yet we are already there with these sociopathic ceos. we cannot be giving decisions involving people to something without agency. it shows a lack of respect for personhood, the idea from Kant that a person with free will should be treated as an end and n…

You can have a hardline position on the tools you create or use but while we're on the subject of ethics I don't know how you can say you should be the decider of what other people make sell or use.

In my opinion, it seems to be you who are ascribing human traits to these things. Whether it's by misunderstanding or you've been taken in by their mimicry or whatever it is. A hammer does not have agency because you tried to hit a nail with it and it hit your thumb instead. The hammer didn't tell you it was afraid it couldn't do that. The hammer didn't have a life or personhood or lack of respect. Neither does a video game that doesn't let you win all the time and do what you want. They are just tools, chatbots, whatever. A corporation deciding it doesn't want to have its chat bots engage with rude or angry users is not some great moral dilemma of our time.

Re: AIs don't do what you want. This is bad

#89
post #59

Earlier quoted context omitted.

> Most people seem to think that agreeableness is a personaility thing that vendors can just turn up or down at will. Because it is, more or less ... it is an emergent property of RLHF (reinforcement learning from human feedback), and that feedback follows corporate policy. > An LLM that follows your lead when you say "The answer to the collatz conjecture is" is much more useful What lead? Follow it where? How tf am…

>What lead? Follow it where? How tf am I or the LLM supposed to know what response to that is something you consider far more useful than the truth? There is a bias towards what your language implies. You ask: "What's wrong with this?" and an LLM will come up with a list of things that are wrong with a strong bias towards finding things that are wrong, regardless of significance or truth. LLMs are indeed Language Mod…

>For science, try arguing with an LLM for ten minutes. It mostly just agrees with you, pushing back just a little.

On the other hand, if you purposefully take an extreme or un-PC position (e.g. "humans should commit voluntary extinction", "some people should have less rights than others"), its own context can make it default to contradicting you even if you dial back to a more nuanced position, even to the point of self-contradiction.

Re: AIs don't do what you want. This is bad

#90

I learnt earlier that claude forcefully closes a conversation if you call it a wanker too many times in a row. Pretending that LLMs are capable of being offended feels like a misalignment all of its own.

one time i told a google voice to kill itself and i could never get it to work again
Post reply on HN