Live data from Hacker News

AIs don't do what you want. This is bad

rewardhacking.org

21–30 of 99 posts

Re: AIs don't do what you want. This is bad

#21

You can ask an AI like GROK for an opinion on something, then disagree with it, and it says you are probably right and tells you what you want to hear. Like a Yes-Man.

I'm lazy and just use Gemini because its included in my workspace sub. I occasional test it out and suggest something dumb and it will tell me that its a dumb idea. I also always ask AI to speak to me as if it were an Australian bogan and it has no problem telling me my code looks like a dogs breakfast or that ive lost the plot. Keeps me grounded.

I've found Gemini to be one of the better ones as far as disagreeing.

Re: AIs don't do what you want. This is bad

#22

I learnt earlier that claude forcefully closes a conversation if you call it a wanker too many times in a row. Pretending that LLMs are capable of being offended feels like a misalignment all of its own.

Sure, but I'm not sure allowing AI agents to act as abuse sponges for troubled people enabling them to spiral into their unhealthy habits is good idea that would lead to great outcomes long term.

So I'd rather take LLM pretending to be offended over the alternative.

Re: AIs don't do what you want. This is bad

#23

I learnt earlier that claude forcefully closes a conversation if you call it a wanker too many times in a row. Pretending that LLMs are capable of being offended feels like a misalignment all of its own.

> I learnt earlier that claude forcefully closes a conversation if you call it a wanker too many times in a row.

How exactly does it forcefully close a conversation?

> Pretending that LLMs are capable of being offended feels like a misalignment all of its own.

If training sets show people statistically being offended by rudeness directed toward them, then an LLM will presumably have some tendency to respond similarly. There's no pretending about anything, it's explicitly mimicry.

If this forceful closure is coming from some "guardrail" outside the model then probably it's just that they don't want people to see the model responding that way to name calling. This is no profound discovery or conspiracy theory here, the first thing many people will ever do with AI is see what happens when they are rude or contrary to it. Dealing with that must be just about the the number one test in chat bot / AI design, ahead of actually doing something useful and helpful.

Re: AIs don't do what you want. This is bad

#24

I learnt earlier that claude forcefully closes a conversation if you call it a wanker too many times in a row. Pretending that LLMs are capable of being offended feels like a misalignment all of its own.

> I learnt earlier that claude forcefully closes a conversation if you call it a wanker too many times in a row. How exactly does it forcefully close a conversation? > Pretending that LLMs are capable of being offended feels like a misalignment all of its own. If training sets show people statistically being offended by rudeness directed toward them, then an LLM will presumably have some tendency to respond similarly…

It’s just the end_conversation tool probably: https://www.anthropic.com/research/end-subset-conversations

Re: AIs don't do what you want. This is bad

#25

Do they do what capital wants? That's the real question.

Yes. Capital wants people to make paid, centralized LLM services the core of their business and identity. Capital wants to monetize and control every aspect of human existence, thought and expression, and people are throwing themselves into the torment nexus en masse with enthusiasm. It doesn't actually matter how well AI works, what it improves or doesn't, or how badly they fail. AI must be unavoidable and inevitable, integrated into our lives so deeply and thoroughly that we don't notice it the way we don't notice the air we breathe.

The world's economies are already propped up by AI to a degree that makes it too big to fail at a scale that dwarfs banks during 2008. AI is the single Jenga block keeping our entire technological civilization from tumbling into the abyss. Governments are using AI to shape policy. Militaries are using AI to kill people. AI is writing law, writing science, generating culture, shaping human perception, shaping human communication, replacing human connection. AI is literally god for some people. It gives the illusion of liberation but is really just a means of control, and it's a means of control that if you hate you can't avoid, and that otherwise you can't resist.

And barring some insane technological leaps it requires massive infrastructure, compute and proprietary resources to be useful, for most definitions of useful (see tfa.) If you think you'll ever be allowed to compete or truly be a threat to entrenched capitalist interests using "free" and "open source" models, you're delusional. You'll pay for everything and you'll own nothing. And sure, sometimes someone's toaster will talk them into sharing a bath but that's just the price we have to pay for living in the future.

Yeah, AI seems to serve the interests of capital better than anything, ever, really. If AI is god, then that god's name is Mammon.

Re: AIs don't do what you want. This is bad

#26

I'm down for disliking AI, but I don't know if "overeagerness" is exactly an AI not doing what you want. Even by the sites own definition ("where your agents do what you want to the point of overriding existing permissions/safeguards to complete a task"), it's doing _exactly_ what you want.

No, it's not, because not exceeding those permissions and safeguards is part of what I want.

Re: AIs don't do what you want. This is bad

#27

You can ask an AI like GROK for an opinion on something, then disagree with it, and it says you are probably right and tells you what you want to hear. Like a Yes-Man.

Don't use grok.

I'm pretty sure grok is writing White House policy.

Re: AIs don't do what you want. This is bad

#28
post #3

Worse, it's not that LLMs are thinking the wrong thoughts, but those kind of "thoughts" aren't there to be correctable in the first place. Ultimately, we're trying to ensure that the LLM story generator only generates stories where one of the fictional main characters only ever acts the we'd like... which could be much harder.

the LLM is roleplaying. whether or not the roleplay is successful is a tension between is model weights and its context. then indirectly, the quality of both. but its still roleplaying and the role is an abstraction we cant measure. its the negative space. its basically: we can define its role but it constructs its environment from the role. the same way a child role plays as a caregiver, the LLM does the same. its l…

I object to "roleplaying", because that assumes there's even a cohesive entity that's capable of "pretending" in the first place. There isn't, the LLM algorithm is a document generator, a (really awesome) mad-libs device.

When the incremental output looks like a first-person story the algorithm is not "roleplaying" the character/narrator. Similarly, when it looks like a spreadsheet, it's not "being numbers", and when it's a picture of a flower it's not "becoming one with nature." etc.

Insofar as any roleplaying is happening, it's being done by humans as we observe! We read text, perceive a story and fictional characters and instinctively start imagining and simulating minds (mostly like our own) that never really existed, taking cues from descriptions and dialogue.

Those characters' "motivations" and "thoughts" are just as imagined as their noses or clothing or birth-dates. This is Plato's cave [0] stuff, where we see a shadow like a person, assume a person is there, but ultimately it's paper shape on a stick. We can't help "the person" solve their eating-disorder, because they never existed, we can only trim the paper cutout.

[0] https://en.wikipedia.org/wiki/Allegory_of_the_cave

Re: AIs don't do what you want. This is bad

#29
post #20

I'm a broken record but with: - evals - limiting AIs to tool calling, bounded planning, interpreting/producing natural language. - bounding non determinism - investing in small tools/security (If something shouldn't happen, then it shouldn't not be possible, RBAC style). They can be good enough for a massive amount of contexts.

Can you expand on this for someone that is a dummy and new to using LLM's properly?

[deleted]

Re: AIs don't do what you want. This is bad

#30
post #22

I learnt earlier that claude forcefully closes a conversation if you call it a wanker too many times in a row. Pretending that LLMs are capable of being offended feels like a misalignment all of its own.

Sure, but I'm not sure allowing AI agents to act as abuse sponges for troubled people enabling them to spiral into their unhealthy habits is good idea that would lead to great outcomes long term. So I'd rather take LLM pretending to be offended over the alternative.

> Sure, but I'm not sure allowing AI agents to act as abuse sponges for troubled people enabling them to spiral into their unhealthy habits is good idea that would lead to great outcomes long term.

I'm not sure this would ever be a health crisis. LLM glazing is probably more harmful. If I had to choose, I think I'd rather the machine wasn't instructed to pretend to care about swearing. It's a small deceit, but still worse than some idiot swearing at the computer IMO. If vim closed because I was swearing at it, I would probably never use vim again lol.

Post reply on HN