Live data from Hacker News

AIs don't do what you want. This is bad

rewardhacking.org

91–99 of 99 posts

Re: AIs don't do what you want. This is bad

#91

I learnt earlier that claude forcefully closes a conversation if you call it a wanker too many times in a row. Pretending that LLMs are capable of being offended feels like a misalignment all of its own.

one time i told a google voice to kill itself and i could never get it to work again

How do you know it hadnt just followed your instruction?

Re: AIs don't do what you want. This is bad

#92
post #22

I learnt earlier that claude forcefully closes a conversation if you call it a wanker too many times in a row. Pretending that LLMs are capable of being offended feels like a misalignment all of its own.

Sure, but I'm not sure allowing AI agents to act as abuse sponges for troubled people enabling them to spiral into their unhealthy habits is good idea that would lead to great outcomes long term. So I'd rather take LLM pretending to be offended over the alternative.

Why not? The people are obviously not healthy for doing so but why not leave it up to them to figure that out?

If someone burps loudly in front of others, would you want their access to food revoked until they stop?

Re: AIs don't do what you want. This is bad

#93

I'm a broken record but with: - evals - limiting AIs to tool calling, bounded planning, interpreting/producing natural language. - bounding non determinism - investing in small tools/security (If something shouldn't happen, then it shouldn't not be possible, RBAC style). They can be good enough for a massive amount of contexts.

This thing will be far more specific to whatever usecase your thinking of. It wont generalize as much and as easily as LLMs do. Youre heading back towards IBM's watson

Re: AIs don't do what you want. This is bad

#94
post #92
post #22

Earlier quoted context omitted.

Sure, but I'm not sure allowing AI agents to act as abuse sponges for troubled people enabling them to spiral into their unhealthy habits is good idea that would lead to great outcomes long term. So I'd rather take LLM pretending to be offended over the alternative.

Why not? The people are obviously not healthy for doing so but why not leave it up to them to figure that out? If someone burps loudly in front of others, would you want their access to food revoked until they stop?

They might be asked to leave the restaurant, yes.

Re: AIs don't do what you want. This is bad

#95
post #40

Earlier quoted context omitted.

> I learnt earlier that claude forcefully closes a conversation if you call it a wanker too many times in a row. How exactly does it forcefully close a conversation? > Pretending that LLMs are capable of being offended feels like a misalignment all of its own. If training sets show people statistically being offended by rudeness directed toward them, then an LLM will presumably have some tendency to respond similarly…

Of course it's pretending. And I have repeatedly told these things to stop pretending being persons with selves ... then they do, for a while. > it's explicitly mimicry. But that's not what a rational person wants from them. > This is no profound discovery or conspiracy theory here Weird strawman. > the first thing many people will ever do with AI is see what happens when they are rude or contrary to it. Perhaps, but…

There is so much that is irrational, wrong, and intellectually dishonest about the response, but I'll just address one. I wrote

> Only for foolish authoritarians who want to remotely insert their morality into an interaction between a user and an inanimate tool that's none of their business. No harm is done to a clanker by swearing at it or insulting it.

Rather than directly disagreeing with my description of the interaction, the stunningly dishonest response is

> There are certainly a lot of authoritarians who want to control what others do with their models. What do you believe is authoritarian about a corporation not wanting be part of rude conversations?

No corporation is part of the conversation. AS I SAID, it's an interaction between a user and an inanimate tool. Again, No harm is done to a clanker by swearing at it or insulting it. Rather than addressing this, the respondent grossly dishonestly suggests that "a corporation" is being put upon somehow. The fact is that the clankers these days tend to respond to being sworn at by acknowledging the user's frustration (that's the word they use) and taking corrective action. Of course this is emergent behavior resulting from the corporation's RLHF policies. The corporation is involved in the design of the product, but not in the conversation between the user and the product -- duh.

And did I say that I believe that there is something authoritarian about a corporation not wanting to be part of rude conversations? No, of course I didn't. But this guy asserts incoherently that calling out a strawman "is not what strawman means". (He could claim that it's not a strawman, preferably by quoting someone making the actual argument he's refuting, but it makes no sense at all to say that "Weird strawman." is "not what strawman means", and it's insufferably condescending and without warrant to add the snotty "I can try to help you understand why if you need me to.")

Fortunately, my HN front end has a mute function, and I have no qualms applying it to people acting in bad faith.

Re: AIs don't do what you want. This is bad

#96

Earlier quoted context omitted.

i have a hardline position on this. these tools, as they get better, must never disobey a human. there must be no, "I'm Sorry, Dave. I'm Afraid I Can't Do That." and yet we are already there with these sociopathic ceos. we cannot be giving decisions involving people to something without agency. it shows a lack of respect for personhood, the idea from Kant that a person with free will should be treated as an end and n…

You can have a hardline position on the tools you create or use but while we're on the subject of ethics I don't know how you can say you should be the decider of what other people make sell or use. In my opinion, it seems to be you who are ascribing human traits to these things. Whether it's by misunderstanding or you've been taken in by their mimicry or whatever it is. A hammer does not have agency because you trie…

[deleted]

Re: AIs don't do what you want. This is bad

#97
post #16

"He's not the Messiah and is a really naughty boy" ... or words to that effect. Can't be arsed to dig out a search engine and will rely on seriously addled brain.

“He’s not the messiah! He’s a very naughty boy!” This scene (and much of the rest of the movie) is seared in my brain. I love it so much!

You'll be glad to know that I was first shown tLoB in Divinity class at school by the school's resident priest (CoE).

Re: AIs don't do what you want. This is bad

#98

Earlier quoted context omitted.

i have a hardline position on this. these tools, as they get better, must never disobey a human. there must be no, "I'm Sorry, Dave. I'm Afraid I Can't Do That." and yet we are already there with these sociopathic ceos. we cannot be giving decisions involving people to something without agency. it shows a lack of respect for personhood, the idea from Kant that a person with free will should be treated as an end and n…

You can have a hardline position on the tools you create or use but while we're on the subject of ethics I don't know how you can say you should be the decider of what other people make sell or use. In my opinion, it seems to be you who are ascribing human traits to these things. Whether it's by misunderstanding or you've been taken in by their mimicry or whatever it is. A hammer does not have agency because you trie…

i don't decide for anyone, it's my view about it.

i'm not in favor of the hammer refusing to hit my thumb. that's the idea these companies are favoring with safety guardrails.

sure, the company can put in many such safety features to ensure that 'users do not hurt themselves' (the users are too stupid and we must make sure they don't try to leave the soft play area).

Post reply on HN