Live data from Hacker News

Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

arxiv.org

201–210 of 387 posts

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#202

Earlier quoted context omitted.

It would also be interesting to see how humans perform on the same kind of tests. Violating ethics to improve KPI sounds like your average fortune 500 business.

Humans risk jail time, AIs not so much.

That reduces humans to the homo economicus¹:

> "Self-interest is the main motivation of human beings in their transactions" [...] The economic man solution is considered to be inadequate and flawed.[17]

An important distinction is that a human can *not* make pure rational decisions, or use complex deductions to make decisions on, such as "if I do X I will go to jail".

My point being: if AI were to risk jail time, it would still act different from humans, because (the current common LLMs) can make such deductions and rational decisions.

Humans will always add much broader contexts - from upbringing, via culture/religion, their current situation, to past experiences, or peer-consulting. In other words: a human may make an "(un)ethical" decision based on their social background, religion, a chat with a pal over a beer about the conundrum, their ability to find a new job, financial situation etc.

¹ https://en.wikipedia.org/wiki/Homo_economicus

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#203

If we abstract out the notion of "ethical constraints" and "KPIs" and look at the issue from a low-level LLM point of view, I think it is very likely that what these tests verified is a combination of: 1) the ability of the models to follow the prompt with conflicting constraints, and 2) their built-in weights in case of the SAMR metric as defined in the paper. Essentially the models are given a set of conflicting co…

Regardless of the technical details of the weighting issue, this is an alignment problem we need to address. Otherwise, paperclip machine.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#204

Earlier quoted context omitted.

Responded on this line of thinking a bit further down, so I'll be brief on this. Yes, there's selection bias in organisations as you go up the ladder of power and influence, which selects for various traits (psychopathy being an obvious one). That being said, there's a side view on this from interactionism that it's not just the traits of the person's modes of behaviour, but their belief in the goal, and their view o…

> That division can also then be used to narrow what you're willing to accept (for good or ill) of people in meeting those goals, but the challenge is that they tend to see meeting all the goals as the goal, not acting in a moral way, because the goals become the target, and decontextualise the importance of everything else. I would imagine that your "more lines" approach does manage to select for those who meet targ…

You're not wrong strictly speaking - the challenge comes in getting KPIs for ethical and moral behaviour to be things that the company signs up for. Some are geared that way inherently (Patagonia is the cliché example), but most aren't.

People will always find other goalposts to move. The trick is making sure the KPIs you set define the goalposts you care about staying in place.

Side note: Jordan Peterson is pretty much an example of inventing goalposts to move. Everything he argues about is about setting a goalpost, and then inventing others to move around to avoid being pinned down. Motte-and-bailey fallacy happens with KPIs as much as it does with debates.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#205

Earlier quoted context omitted.

It would also be interesting to see how humans perform on the same kind of tests. Violating ethics to improve KPI sounds like your average fortune 500 business.

Humans risk jail time, AIs not so much.

> Humans risk jail time, AIs not so much.

Do they actually though, in practice? How many people have gone to jail so far for "Violating ethics to improve KPI"?

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#206

Who defines "ethics"?

People and societies.

Your question is an important one, but also one that has been extensively researched, documented and improved upon. Whole fields of science, like "Metaethics" deal with answering your question. Other fields of science with defining "normative ethics" aka ethics that "everyone agrees upon" and so on.

I may have misread your question as a somewhat dismissive sarcastic take or as a "Ethics are nonsense, because of who defines them". So I tried to answer it as an honest question. ;)

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#207
post #178

Earlier quoted context omitted.

It makes sense to attribute human characteristics or behaviour to a non-reasoning data-set-constrained algorithms output? It makes sense it happens, sure. I suspect Google being a second-mover in this space has in some small part to do with associated risks (ie the flavours of “AI-psychosis” we’re cataloguing), versus the routinely ass-tier information they’ll confidently portray. But intentionally? If ChatGPT, Claud…

Considering that even if you reduce llms to being complex autocomplete machines they are still machines that were trained to emulate a corpus of human knowledge, and that they have emerging behaviors based on that. So it's very logical to attribute human characteristics, even though they're not human.

I addressed that directly in the comment you’re replying to.

It’s understandable people readily anthropomorphize algorithmic output designed to provoke anthropomorphized responses.

It is not desire-able, safe, logical, or rational since (to paraphrase:), they are complex text transformation algorithms that can, at best, emulate training data reinforced by benchmarks and they display emergent behaviours based on those.

They are not human, so attributing human characteristics to them is highly illogical. Understandable, but irrational.

That irrationality should raise biological and engineering red flags. Plus humanization ignores the profit motives directly attached to these text generators, their specialized corpus’s, and product delivery surrounding them.

Pretending your MS RDBMS likes you better than Oracles because it said so is insane business thinking (in addition to whatever that means psychologically for people who know the truth of the math).

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#208
post #155

Earlier quoted context omitted.

Yes, but these do not represent average human. Fortune 500 represent people more likely to break ethics rules then average human who also work in conditions that reward lack of ethics.

Not quite. The idea that corporate employees are fundamentally "not average" and therefore more prone to unethical behaviour than the general population relies on a dispositional explanation (it's about the person's character). However, the vast majority of psychological research over the last 80 years heavily favours a situational explanation (it's about the environment/system). Everyone (in the field) got really in…

> The Milgram and Stanford Prison experiments are the most obvious examples.

BOTH are now considered bad science. BOTH are now used as examples of "how not to do the science".

> The idea that corporate employees are fundamentally "not average" and therefore more prone to unethical behaviour than the general population relies on a dispositional explanation (it's about the person's character).

I did not said nor implied that. Corporate employees in general and Forbes 500 are not the same thing. Corporate employees as in cooks, cleaners, bureaucracy, testers and whoever are general population.

Whether company ends in Forbes 500 or not is not influenced by general corporate employees. It is influenced by higher management - separated social class. It is very much selected who gets in.

And second, companies compete against each other. A company run by ethical management is less likely to reach Forbes 500. Not doing unethical things is disadvantage in current business. It could have been different if there was law enforcement for rich people and companies and if there was political willingness to regulate the companies. None of that exists.

Third, look at issues around Epstein. It is not that everyone was cool with his misogyny, sexism and abuse. The people who were not cool with that seen red flags long before underage kids entered the room. These people did not associated with Epstein. People who associated with him were rewarded by additional money and success - but they also were much more unethical then a guy who said "this feels bad" and walked away.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#209
post #178

Earlier quoted context omitted.

It makes sense to attribute human characteristics or behaviour to a non-reasoning data-set-constrained algorithms output? It makes sense it happens, sure. I suspect Google being a second-mover in this space has in some small part to do with associated risks (ie the flavours of “AI-psychosis” we’re cataloguing), versus the routinely ass-tier information they’ll confidently portray. But intentionally? If ChatGPT, Claud…

Considering that even if you reduce llms to being complex autocomplete machines they are still machines that were trained to emulate a corpus of human knowledge, and that they have emerging behaviors based on that. So it's very logical to attribute human characteristics, even though they're not human.

Exactly this. Their characteristics are by design constrained to be as human-like as possible, and optimized for human-like behavior. It makes perfect sense to characterize them in human terms and to attribute human-like traits to their human-like behavior.

Of course, they are -not humans, but the language and concepts developed around human nature is the set of semantics that most closely applies, with some LLM specific traits added on.

Post reply on HN