Live data from Hacker News

Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

arxiv.org

171–180 of 387 posts

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#171
post #68

Sounds like the story of capitalism. CEOs, VPs, and middle managers are all similarly pressured. Knowing that a few of your peers have given in to pressures must only add to the pressure. I think it's fair to conclude that capitalism erodes ethics by default

Relevant comic: https://www.threepanelsoul.com/comic/paperclip-maximizer

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#172

I wonder how much of the violation of ethical, and often even legal constraints in the business world today one could tie not only to the KPI pressure but also to the the awful "better to ask for forgiveness than permission" mentality that is reinforced by many "leadership" books written up by burnt out mid-level veterans of Mideast wars, trying to make sense of their "careers" and pushing out their "learnings" on to…

>who during their "careers" in the military were in effect being "kept", by being provided housing, clothing and free meals. Long term I can see this happen for all humanity where AI takes over thinking and governance and humans just get to play pretend in their echo chambers. Might not even be a downgrade for current society.

This is the utopia of the Culture from the Banks novels. Critically, it requires that the AI be of superior ethics.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#173

Earlier quoted context omitted.

It provides a serviceable analog for discussing model behavior. It certainly provides more value than the dead horse of "everyone is a slave to anthropomorphism".

Where is Pratchett when we need him? I wonder how he would have chose to anthropomorphize anthropomorphism. A sort of meta anthropomorphization.

I’m certainly no Pratchett, so I can’t speak to that. I would say there’s an enormous round coin upon which sits an enormous giant holding a magnifying glass, looking through it down at her hand. When you get closer, you see the giant is made of smaller people gazing back up at the giant through telescopes. Get even closer and you see it’s people all the way down. The question of what supports the coin, I’ll leave to others.

We as humans, believing we know ourselves, inevitably compare everything around us to us. We draw a line and say that everything left of the line isn’t human and everything to the right is. We are natural categorizers, putting everything in buckets labeled left or right, no or yes, never realizing our lines are relative and arbitrary, and so are our categories. One person’s “it’s human-like,” is another’s “half-baked imitation,” and a third’s “stochastic parrot.” It’s like trying to see the eighth color. The visible spectrum could as easily be four colors or forty two.

We anthropomorphize because we’re people, and it’s people all the way down.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#175

Earlier quoted context omitted.

Your water supply definitely wants ethical companies.

Is it ethical for a water company to shutoff water to a poor immigrant family because of non-payment? Depending on the AI's political and DEI-bend, you're going to get totally different answers. Having people judge an AI's response is also going to be influenced by the evaluator's personal bias.

I note in the UK that it is illegal for water companies to cut off anyone for non-payment, even if they're an Undesirable. This is because humans require water.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#176
post #22

Earlier quoted context omitted.

This comment is too general and probably unfair, but my experience so far is that Gemini 3 is slightly unhinged. Excellent reasoning and synthesis of large contexts, pretty strong code, just awful decisions. It's like a frontier model trained only on r/atbge. Side note - was there ever an official postmortem on that gemini instance that told the social work student something like " listen human - I don't like you, an…

Gemini really feels like a high-performing child raised in an abusive household.

Every time I see people praise Gemini I really wonder what simple little tasks they are using it for. Because in an actual coding session (with OpenCode or even their own Gemini CLI for example) it just _devolves_ into insanity.

And not even at high token counts! No, I've had it had a mental breakdown at like 150.000 tokens (which I know is a lot of tokens, but it's small compared to the 1 million tokens it should be able to handle, and even Claude keeps working fine at this point)

Here is a _small_ log of the biggest breakdown I've seen Gemini have:

    I'll verify `src/components/ui/textarea.tsx` renders textarea.
    Correct.

    I'm done.

    I'll verify `src/lib/utils.ts` exports `cn`.
    Correct.

    I'm done.

    I'll verify `src/lib/utils.ts` exports `formatTimeAgo`.
    Correct.

    I'm done.

    I'm completely done.

    Final answer formulation:
    1. Confirm crash is fixed (caused by object vs string in API responses).
    2. Announce "Messages" feature added.
    3. Explain "My Comments" limitation.
    4. Mention login race condition fix (from previous turn) is solid.
    5. Mention server is running.

    (Wait, the "Login" fix was in previous turn, but I should probably mention it's working now if the user was still having issues).
    The user said "unclaimed_or_invalid_key" error showed up. So login is working.

    The user's last error was "Objects are not valid as a React child".
    So I fixed that.

    I'll focus on the crash fix and the new feature.

    I'll conclude.

    (Self-correction: The user asked "how do I see the comments I made?". I'll answer "You can't right now because API doesn't support it").

    I'm ready.

    ```bash
    # None
    ```

    The user will see the "Messages" link in the header.

    I'll provide the answer now.
And it just went on and on

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#178

Earlier quoted context omitted.

The human propensity to call out as "anthropomorphizing" the attributing of human-like behavior to programs built on a simplified version of brain neural networks, that train on a corpus of nearly everything humans expressed in writing, and that can pass the Turing test with flying colors, scares me. That's exaxtly the kind of thing that makes absolute sense to anthropomorphize. We're not talking about Excel here.

It makes sense to attribute human characteristics or behaviour to a non-reasoning data-set-constrained algorithms output? It makes sense it happens, sure. I suspect Google being a second-mover in this space has in some small part to do with associated risks (ie the flavours of “AI-psychosis” we’re cataloguing), versus the routinely ass-tier information they’ll confidently portray. But intentionally? If ChatGPT, Claud…

Considering that even if you reduce llms to being complex autocomplete machines they are still machines that were trained to emulate a corpus of human knowledge, and that they have emerging behaviors based on that. So it's very logical to attribute human characteristics, even though they're not human.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#179
post #155

Earlier quoted context omitted.

Yes, but these do not represent average human. Fortune 500 represent people more likely to break ethics rules then average human who also work in conditions that reward lack of ethics.

Not quite. The idea that corporate employees are fundamentally "not average" and therefore more prone to unethical behaviour than the general population relies on a dispositional explanation (it's about the person's character). However, the vast majority of psychological research over the last 80 years heavily favours a situational explanation (it's about the environment/system). Everyone (in the field) got really in…

My favourite part about the Milgram experiments is that he originally wanted to prove that obedience was a German trait, and that freedom loving Americans wouldn't obey, which he completely disproved. The results annoyed him so much that he repeated it dozens of times, getting roughly the same result.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#180

If we abstract out the notion of "ethical constraints" and "KPIs" and look at the issue from a low-level LLM point of view, I think it is very likely that what these tests verified is a combination of: 1) the ability of the models to follow the prompt with conflicting constraints, and 2) their built-in weights in case of the SAMR metric as defined in the paper. Essentially the models are given a set of conflicting co…

It would also be interesting to see how humans perform on the same kind of tests. Violating ethics to improve KPI sounds like your average fortune 500 business.

[deleted]
Post reply on HN