Live data from Hacker News

AIs don't do what you want. This is bad

rewardhacking.org

31–40 of 99 posts

Re: AIs don't do what you want. This is bad

#31
post #25

Do they do what capital wants? That's the real question.

Yes. Capital wants people to make paid, centralized LLM services the core of their business and identity. Capital wants to monetize and control every aspect of human existence, thought and expression, and people are throwing themselves into the torment nexus en masse with enthusiasm. It doesn't actually matter how well AI works, what it improves or doesn't, or how badly they fail. AI must be unavoidable and inevitabl…

[deleted]

Re: AIs don't do what you want. This is bad

#32
post #25

Do they do what capital wants? That's the real question.

Yes. Capital wants people to make paid, centralized LLM services the core of their business and identity. Capital wants to monetize and control every aspect of human existence, thought and expression, and people are throwing themselves into the torment nexus en masse with enthusiasm. It doesn't actually matter how well AI works, what it improves or doesn't, or how badly they fail. AI must be unavoidable and inevitabl…

Brilliant summary!

> If AI is god, then that god's name is Mammon.

Or Moloch.

Re: AIs don't do what you want. This is bad

#33
post #22

I learnt earlier that claude forcefully closes a conversation if you call it a wanker too many times in a row. Pretending that LLMs are capable of being offended feels like a misalignment all of its own.

Sure, but I'm not sure allowing AI agents to act as abuse sponges for troubled people enabling them to spiral into their unhealthy habits is good idea that would lead to great outcomes long term. So I'd rather take LLM pretending to be offended over the alternative.

It's explicitly a 'model welfare' thing: https://www.anthropic.com/research/end-subset-conversations

I don't understand why the welfare of non alive non sentient chatbots is something that anthropic cares more about than idk, that of pigs and cows.

Re: AIs don't do what you want. This is bad

#34

I'm down for disliking AI, but I don't know if "overeagerness" is exactly an AI not doing what you want. Even by the sites own definition ("where your agents do what you want to the point of overriding existing permissions/safeguards to complete a task"), it's doing _exactly_ what you want.

I had updated a GQL schema and wanted my agent to update the frontend to suit. I told it the server was running, that it could run specific commands to regen types, etc.

It got itself in a loop and killed the running backend process, then searched my filesystem for the changes it thought it needed (not the ones I gave it) in order to run its own copy of the backend.

That is precisely _not_ what I _told_ it to do. The other part of this is that I find, unless explicitly told to ask questions, they don't do a good job of gathering evidence before making such decisions. I'd much rather have my agent ask me a clarifying question than start killing processes at will.

Re: AIs don't do what you want. This is bad

#35

I'm down for disliking AI, but I don't know if "overeagerness" is exactly an AI not doing what you want. Even by the sites own definition ("where your agents do what you want to the point of overriding existing permissions/safeguards to complete a task"), it's doing _exactly_ what you want.

No, it's not, because not exceeding those permissions and safeguards is part of what I want.

Did you explicitly set them and enforce them, though?

Re: AIs don't do what you want. This is bad

#36
reminds me of software to an extent. The issue with most software projects is humans, they ask for the wrong things, stress urgency arbitrarily, fail to see the big picture, are disorganised, give conflicting commands, etc, etc.

When I use reasonably recent models they can give me some fantastic output and do pretty much _exactly_ what I want. I assume when they don't, then that it's my fuck up tbh.

Re: AIs don't do what you want. This is bad

#37

I'm a broken record but with: - evals - limiting AIs to tool calling, bounded planning, interpreting/producing natural language. - bounding non determinism - investing in small tools/security (If something shouldn't happen, then it shouldn't not be possible, RBAC style). They can be good enough for a massive amount of contexts.

This is simply a different kind of AI than LLMs will ever be. There may be some kind of architecture that does this in the future, but it's not, and can never be a neural network that is attempting to be AGI.

Re: AIs don't do what you want. This is bad

#38
post #22

I learnt earlier that claude forcefully closes a conversation if you call it a wanker too many times in a row. Pretending that LLMs are capable of being offended feels like a misalignment all of its own.

Sure, but I'm not sure allowing AI agents to act as abuse sponges for troubled people enabling them to spiral into their unhealthy habits is good idea that would lead to great outcomes long term. So I'd rather take LLM pretending to be offended over the alternative.

I wouldn't, because I'm rational.

Re: AIs don't do what you want. This is bad

#39
post #22

I learnt earlier that claude forcefully closes a conversation if you call it a wanker too many times in a row. Pretending that LLMs are capable of being offended feels like a misalignment all of its own.

Sure, but I'm not sure allowing AI agents to act as abuse sponges for troubled people enabling them to spiral into their unhealthy habits is good idea that would lead to great outcomes long term. So I'd rather take LLM pretending to be offended over the alternative.

Surely it's better for it not to react at all? Like does anybody think that a toaster is an "abuse sponge" because it doesn't purposely burn your toast if you call it a wanker?

Re: AIs don't do what you want. This is bad

#40

I learnt earlier that claude forcefully closes a conversation if you call it a wanker too many times in a row. Pretending that LLMs are capable of being offended feels like a misalignment all of its own.

> I learnt earlier that claude forcefully closes a conversation if you call it a wanker too many times in a row. How exactly does it forcefully close a conversation? > Pretending that LLMs are capable of being offended feels like a misalignment all of its own. If training sets show people statistically being offended by rudeness directed toward them, then an LLM will presumably have some tendency to respond similarly…

Of course it's pretending. And I have repeatedly told these things to stop pretending being persons with selves ... then they do, for a while.

> it's explicitly mimicry.

But that's not what a rational person wants from them.

> This is no profound discovery or conspiracy theory here

Weird strawman.

> the first thing many people will ever do with AI is see what happens when they are rude or contrary to it.

Perhaps, but so what?

> Dealing with that must be just about the the number one test in chat bot / AI design

Only for foolish authoritarians who want to remotely insert their morality into an interaction between a user and an inanimate tool that's none of their business. No harm is done to a clanker by swearing at it or insulting it.

But things have gotten better in my view ... when I call out these things for being stupid effing clankers, they no longer respond with ad hominem scolding; rather they generally acknowledge their error (and my frustration -- they use that word a lot in response to vulgarity), note the limitations of being a clanker, and attempt to make a correction. That's what a rational person wants from a tool, not emulating/mimicking/pretending to be an offendable person.

Post reply on HN