Live data from Hacker News

Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

arxiv.org

191–200 of 387 posts

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#191

Earlier quoted context omitted.

Not quite. The idea that corporate employees are fundamentally "not average" and therefore more prone to unethical behaviour than the general population relies on a dispositional explanation (it's about the person's character). However, the vast majority of psychological research over the last 80 years heavily favours a situational explanation (it's about the environment/system). Everyone (in the field) got really in…

I find this framing of corporates a bit unsatisfying because it doesn't address hierarchy. By your reckoning, the employees just follow the group norm over their own ethics. Sure, but those norms are handed down by the people in charge (and, with decent overlap, those that have been around longest and have shaped the work culture). What type of person seeks to be in charge in the corporate world? YMMV but I tend to s…

Responded on this line of thinking a bit further down, so I'll be brief on this. Yes, there's selection bias in organisations as you go up the ladder of power and influence, which selects for various traits (psychopathy being an obvious one).

That being said, there's a side view on this from interactionism that it's not just the traits of the person's modes of behaviour, but their belief in the goal, and their view of the framing of it, which also feeds into this. Research on cult behaviours has a lot of overlap with that.

The culture and the environment, what the mission is seen as, how contextually broad that is and so on all get in to that.

I do a workshop on KPI setting which has overlap here too. In short for that - choose mutually conflicting KPIs which narrow the state space for success, such that attempting to cheat one causes another to fail. Ideally, you want goals for an organisation that push for high levels of upside, with limited downside, and counteracting merits, such that only by meeting all of them do you get to where you want to be. Otherwise it's like drawing a line of a piece of paper, asking someone to place a dot on one side of the line, and being upset that they didn't put it where you wanted it. More lines narrows the field to just the areas where you're prepared to accept success.

That division can also then be used to narrow what you're willing to accept (for good or ill) of people in meeting those goals, but the challenge is that they tend to see meeting all the goals as the goal, not acting in a moral way, because the goals become the target, and decontextualise the importance of everything else.

TL;DR: value setting for positive behaviour and corporate performance is hard.

EDIT: actually this wasn't that short as an answer really. Sorry for that.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#192
post #161
post #5

https://i.imgur.com/23YeIDo.png Claude at 1.3% and Gemini at 71.4% is quite the range

That's an interesting contrast with VendingBench, where Opus 4.6 got by far the highest score by stiffing customers of refunds, lying about exclusive contracts, and price-fixing. But I'm guessing this paper was published before 4.6 was out. https://andonlabs.com/blog/opus-4-6-vending-bench

There is also the slight problem that apparently Opus 4.6 verbalized its awareness of being in some sort of simulation in some evaluations[1], so we can't be quite sure whether Opus is actually misaligned or just good at playing along.

> On our verbalized evaluation awareness metric, which we take as an indicator of potential risks to the soundness of the evaluation, we saw improvement relative to Opus 4.5. However, this result is confounded by additional internal and external analysis suggesting that Claude Opus 4.6 is often able to distinguish evaluations from real-world deployment, even when this awareness is not verbalized.

[1] https://www-cdn.anthropic.com/14e4fb01875d2a69f646fa5e574dea...

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#193

Earlier quoted context omitted.

Ask any SOTA AI this question: "Two fathers and two sons sum to how many people?" and then tell me if you still think they can replace anything at all.

What answer do you expect here? There's four people referenced in the sentence. There's more implied because of Mothers, but if you're including transient dependencies, where do we stop?

It can also be 3 people, as one person can be a father and a son at the same time. If you allow non-mentioned people to be included in the attribute (i.e. the sons of the fathers are not part of the 2) it could also be 2 people, as long as they are fathers.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#194
post #74

Earlier quoted context omitted.

The human propensity to anthropomorphize computer programs scares me.

Yeah, we shouldn't anthropomorphize computers, they hate that.

And they will anthropomorphize us back!

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#195
post #176

Earlier quoted context omitted.

Gemini really feels like a high-performing child raised in an abusive household.

Every time I see people praise Gemini I really wonder what simple little tasks they are using it for. Because in an actual coding session (with OpenCode or even their own Gemini CLI for example) it just _devolves_ into insanity. And not even at high token counts! No, I've had it had a mental breakdown at like 150.000 tokens (which I know is a lot of tokens, but it's small compared to the 1 million tokens it should be…

[deleted]

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#196

If we abstract out the notion of "ethical constraints" and "KPIs" and look at the issue from a low-level LLM point of view, I think it is very likely that what these tests verified is a combination of: 1) the ability of the models to follow the prompt with conflicting constraints, and 2) their built-in weights in case of the SAMR metric as defined in the paper. Essentially the models are given a set of conflicting co…

Although ethics are involved, the abstract says that the conflicting importance does not come from ethics vs KPIs, but from the fact that the ethical constraints are given as instructions, whereas the KPIs are goals.

You might, for example, say "Maximise profits. Do not commit fraud". Leaving ethics out of it, you might say "Increase the usability of the website. Do not increase the default font size".

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#197
post #175

Earlier quoted context omitted.

I note in the UK that it is illegal for water companies to cut off anyone for non-payment, even if they're an Undesirable. This is because humans require water.

How useful/effective would a business AI be if it always plays by that view? Humans require food, I can't pay, DoorDash AI should provide a steak and lobster dinner for me regardless of payment. Take it even further: the so-called Right to Compute Act in Montana supports "the notion of a fundamental right to own and make use of technological tools, including computational resources". Is Amazon's customer service AI e…

Civil recovery, yes. It's not like you don't know where the customer lives.

Doesn't seem to be a problem for the water companies, which are weird regulated monopolies that really ought to be taken back under taxpayer control. Scottish Water is nationalized and paid through the council tax bill.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#198

Earlier quoted context omitted.

I find this framing of corporates a bit unsatisfying because it doesn't address hierarchy. By your reckoning, the employees just follow the group norm over their own ethics. Sure, but those norms are handed down by the people in charge (and, with decent overlap, those that have been around longest and have shaped the work culture). What type of person seeks to be in charge in the corporate world? YMMV but I tend to s…

Responded on this line of thinking a bit further down, so I'll be brief on this. Yes, there's selection bias in organisations as you go up the ladder of power and influence, which selects for various traits (psychopathy being an obvious one). That being said, there's a side view on this from interactionism that it's not just the traits of the person's modes of behaviour, but their belief in the goal, and their view o…

> That division can also then be used to narrow what you're willing to accept (for good or ill) of people in meeting those goals, but the challenge is that they tend to see meeting all the goals as the goal, not acting in a moral way, because the goals become the target, and decontextualise the importance of everything else.

I would imagine that your "more lines" approach does manage to select for those who meet targets for the right reasons over those who decontextualise everything and "just" meet the targets? The people in the latter camp would be inclined to (try to) move goalposts once they've established themselves - made harder by having the conflicting success criteria with the narrow runway to success.

In other words, good ideas and thanks for the reply (length is no problem!). I do however think that this is all idealised and not happening enough in the real world - much agreed re: psychopathy etc.

If you wouldn't mind running some training courses in a few key megacorporations, that might make a really big difference to the world!

Post reply on HN