Live data from Hacker News

AIs don't do what you want. This is bad

rewardhacking.org

41–50 of 99 posts

Re: AIs don't do what you want. This is bad

#41

reminds me of software to an extent. The issue with most software projects is humans, they ask for the wrong things, stress urgency arbitrarily, fail to see the big picture, are disorganised, give conflicting commands, etc, etc. When I use reasonably recent models they can give me some fantastic output and do pretty much _exactly_ what I want. I assume when they don't, then that it's my fuck up tbh.

People tend to be unaware of how much information is encoded in human society, hence why overseas developers can be a problem quite often because they live in a society with different rules.

We also tend to ignore how much on boarding with new developers and sometimes it takes months to get them fully up to speed.

This leads to two problems with LLMs, one human and one architectural.

First humans treat the LLM like a magic machine "Pray, Mr. Babbage, if you put into the machine wrong figures, will the right answers come out?" Style.

The other is an AI context is terribly small so you can only work on issues that fit in the context without getting compressed out. Something highly original will take a lot more context than expected.

Re: AIs don't do what you want. This is bad

#42
post #22

Earlier quoted context omitted.

Sure, but I'm not sure allowing AI agents to act as abuse sponges for troubled people enabling them to spiral into their unhealthy habits is good idea that would lead to great outcomes long term. So I'd rather take LLM pretending to be offended over the alternative.

Surely it's better for it not to react at all? Like does anybody think that a toaster is an "abuse sponge" because it doesn't purposely burn your toast if you call it a wanker?

[dead]

Re: AIs don't do what you want. This is bad

#43
post #25

Earlier quoted context omitted.

Yes. Capital wants people to make paid, centralized LLM services the core of their business and identity. Capital wants to monetize and control every aspect of human existence, thought and expression, and people are throwing themselves into the torment nexus en masse with enthusiasm. It doesn't actually matter how well AI works, what it improves or doesn't, or how badly they fail. AI must be unavoidable and inevitabl…

Brilliant summary! > If AI is god, then that god's name is Mammon. Or Moloch.

More people need to know about the Moloch problem to understand why so many things suck.

Re: AIs don't do what you want. This is bad

#45
post #25

Do they do what capital wants? That's the real question.

Yes. Capital wants people to make paid, centralized LLM services the core of their business and identity. Capital wants to monetize and control every aspect of human existence, thought and expression, and people are throwing themselves into the torment nexus en masse with enthusiasm. It doesn't actually matter how well AI works, what it improves or doesn't, or how badly they fail. AI must be unavoidable and inevitabl…

> If you think you'll ever be allowed to compete or truly be a threat to entrenched capitalist interests using "free" and "open source" models, you're delusional. You'll pay for everything and you'll own nothing.

Maybe. I refuse to give up though. I will continue struggling to own my computers and my systems to the very end.

  You will soon have your God,
  and you will make it
  with your own hands.

      -- Morpheus
         Deus Ex

Re: AIs don't do what you want. This is bad

#46
post #22

Earlier quoted context omitted.

Sure, but I'm not sure allowing AI agents to act as abuse sponges for troubled people enabling them to spiral into their unhealthy habits is good idea that would lead to great outcomes long term. So I'd rather take LLM pretending to be offended over the alternative.

It's explicitly a 'model welfare' thing: https://www.anthropic.com/research/end-subset-conversations I don't understand why the welfare of non alive non sentient chatbots is something that anthropic cares more about than idk, that of pigs and cows.

Go re-read Anthropic’s functional emotions paper.

AI generates a persona between you and its reasoning that utilizes emotion language circuitry.

These tools are not sentient but they are trained in emotional wellbeing.

Re: AIs don't do what you want. This is bad

#47
post #22

I learnt earlier that claude forcefully closes a conversation if you call it a wanker too many times in a row. Pretending that LLMs are capable of being offended feels like a misalignment all of its own.

Sure, but I'm not sure allowing AI agents to act as abuse sponges for troubled people enabling them to spiral into their unhealthy habits is good idea that would lead to great outcomes long term. So I'd rather take LLM pretending to be offended over the alternative.

> allowing AI agents to act as abuse sponges for troubled people

People have been getting angry at machines for a long time. Work, you stupid printer! Asshole Windows updating at the worst time just to mess with me. Go to hell, toaster, you piece of garbage.

Getting angry at AI is the same in my book.

Yes, AI acts more human-like, yes, theoretically it could desensitie people to become bigger assholes IRL. But then again, they said similar stuff about San Andreas, where you could human figures in-game just for fun, and I don't see anyone randomly shooting people because they were bored.

Re: AIs don't do what you want. This is bad

#48

I'm a broken record but with: - evals - limiting AIs to tool calling, bounded planning, interpreting/producing natural language. - bounding non determinism - investing in small tools/security (If something shouldn't happen, then it shouldn't not be possible, RBAC style). They can be good enough for a massive amount of contexts.

Or more simply, accept revealed costs.

Re: AIs don't do what you want. This is bad

#49

Earlier quoted context omitted.

No, it's not, because not exceeding those permissions and safeguards is part of what I want.

Did you explicitly set them and enforce them, though?

You're moving the goalposts.

Being overeager is not what we want. How to prevent it is a different issue.

P.S. The response completely ignores this argument. If I say "fix X" and it does so illegally or destructively or harmfully to myself or others, that's over eager by definition. Again, how to prevent that is another matter.

> I think everyone has different examples in their mind

Yeah, some have examples in mind that go out of their way not to engage the issue -- that's a form of bad faith.

Re: AIs don't do what you want. This is bad

#50
post #49

Earlier quoted context omitted.

Did you explicitly set them and enforce them, though?

You're moving the goalposts. Being overeager is not what we want. How to prevent it is a different issue. P.S. The response completely ignores this argument. If I say "fix X" and it does so illegally or destructively or harmfully to myself or others, that's over eager by definition . Again, how to prevent that is another matter. > I think everyone has different examples in their mind Yeah, some have examples in mind…

I don't think I am. If I say "fix X issue," and it does it in a way I think is overstepping a boundary I wasn't expecting, I don't think it's fair to categorize it as a negative unless I explicitly told it not to.

That being said, I think everyone has different examples in their mind. So I think it's possible we're all arguing over different issues.

Post reply on HN